Memo "Sarashina" is a large-scale language model (LLM) originating in Japan, developed and provided by SB Intuitions (a SoftBank subsidiary).
This is expected to promote investment in domestic AI technology and strengthen the technology and human resources base that contributes to GDP.
Key features include:
Various model sizes: Sarashina1 (7 billion, 13 billion, and 65 billion parameters) and Sarashina2 (7 billion and 13 billion) models are publicly available (ITmedia, Hi-Res Inc. | GPU Cloud Services).
High performance: On a standard Japanese evaluation dataset, it outperforms other models of the same size. Sarashina2-13B in particular received high praise (SB Intuitions Inc.).
Lightweight Model Development: The Sarashina 2.2 series (500 million to 3 billion parameters) and the 3B model are light enough to run on home PCs (Ledge.ai).
Mixture of Experts Configuration: Sarashina 2-8x70B combines multiple expert models to achieve advanced inference capabilities and is currently available for research purposes (Rozetta Square). Furthermore, the 70 billion parameter "Sarashina mini" is also planned for commercial release (BUSINESS NETWORK).
Understood. Below, we provide a more detailed and systematic introduction to Sarashina, a generative AI model specialized for Japanese.
Sarashina Overview and Model Variations
Developer: A large-scale language model (LLM) specialized for Japanese, developed by SB Intuitions, a subsidiary of SoftBank (Hi-Res Corporation | GPU Cloud Services).
Origin of the name: "Sarashina Diary" (Sarashina Diary), derived from Takeshiba, the location of SoftBank's headquarters (SB Intuitions, Inc.).
Model Configuration and Evolution
Sarashina1 Series (Pre-trained Model): 70B, 130B, and 650B parameter counts. Trained exclusively on Japanese, learning approximately 1 trillion tokens (Hi-Res, Inc. | GPU Cloud Services).
Sarashina2 Series: 70B and 130B models, training over 2.1 trillion tokens including Japanese, English, and programming code (ratio: Japanese 5: English 4: Code 1) (Hi-Res, Inc. | GPU Cloud Services).
Sarashina2-8x70B: Mixture of Experts (MoE) configuration combines eight expert models to demonstrate advanced inference capabilities (total parameter count equivalent to approximately 4,600B) (Rozetta Square).
Learning Environment and Development Platform
AI Computing Infrastructure: Utilizing a supercomputer platform with approximately 6,000 high-performance GPUs, including NVIDIA Hopper and H100 GPUs (ITmedia).
Expansion Plan: Plans to expand the GPU infrastructure to 10,000 units by the first half of fiscal year 2025 (Rozetta Square).
Model Performance and Application Deployment
Natural Language Understanding: Achieves high average performance on multiple Japanese evaluation datasets, including JCommonsenseQA, JEMHopQA, NIILC-QA, JSQuAD, and AI-O (SB Intuitions, Inc.).
Vision-Enabled Model: The Sarashina2-Vision series (8B and 14B) is a vision and language model (VLM) specialized for visual information and Japanese language understanding, achieving robust performance even with Japanese culture, customs, and complex diagrams. Commercial use permitted (MIT License) (SB Intuitions, Inc.).
Lightweight and Embedded Models
Sarashina2.1-1B: A 1B parameter model. First, it was pre-trained on a large-scale Japanese and English corpus (10 trillion tokens), then fine-tuned on a Japanese-only corpus (1 trillion tokens). It achieved solid performance in both Japanese and English language tasks (Hugging Face).
Sarashina-Embedding-v1-1B: A sentence embedding model built on Sarashina2.1-1B. It supports up to 8,192 tokens and generates 1,792-dimensional vectors. It achieved the best average performance across 16 datasets on the JMTEB (Japanese Massive Text Embedding Benchmark) (Hugging Face).
Sarashina2.2 Small Model: Lightweight series with 500, 1, and 3 billion parameters are available under the MIT license and are compatible with a wide range of devices (Ledge.ai).
Sarashina2.2-3B Instruct (GGUF format): A 3B model tuned specifically for command responses. The GGUF format allows for highly efficient local deployment. It also works well with LLAMA.cpp (promptlayer.com).
Outlook for Productization and Social Implementation
Preparing for commercial service: The product team is currently considering specifications and ease of use with an eye toward adoption by corporate users and others. The service is scheduled to begin in fiscal year 2026 (next year) (news.mynavi.jp).
These are the details of "Sarashina" that have been revealed so far.
Digging deeper from a GDP perspective, Sarashina will promote "in-house AI development" centered on SB Intuitions' Japanese LLMs (7B/13B/65B/70B, 1B series, Embedding/Vision series). This will simultaneously curb the outflow of overseas model usage fees and attract domestic GPU investment (large-scale clusters on the H100/Blackwell scale), strengthening the foundation of real GDP through deepening ICT capital and boosting TFP. SB Intuitions Inc. ITmedia RCR Wireless News
In the field, automation of internal search, summarization, and call center functions will boost labor productivity in the service industry. Specializing in Japanese language reduces implementation costs, and compliance with "sovereign AI" requirements (data sovereignty for public and medical services) will promote procurement. This is expected to have a multiplier effect: capital investment → demand for electricity and data centers → related employment. SoftBank ITmedia
Distribution will proceed via Hugging Face (e.g., Sarashina 2.1-1B, Embedded 1B, Vision 14B). Non-commercial licenses and the availability of Japanese language benchmarks facilitate the transition from PoC to production. The final GDP contribution will depend largely on the ratio of in-house development to adoption rate.
Domestic supply of the Sarashina series (Sarashina 1/2, Sarashina 2-Vision, 2.1-1B, and 2.2 SLM) will enhance the domestic circulation of IT spending through import substitution for inference/APIs and easier adoption of domestic vendors, strengthening the GDP path of improved TFP and boosting potential growth. (SB Intuitions, Inc., Hugging Face, Ledge.ai)
In terms of licensing, for example, Sarashina 1-65B is MIT-approved for easy redistribution and commercial use, 2.1-1B is non-commercial, and 2.2 SLM is commercially available, allowing for different uses. In-house development of RAG, summarization, and dialogue can increase added value and time in the service industry. (Hugging Face, Ledge.ai, SoftBank)
Furthermore, the MoE's Sarashina 2-8x70B and the "Sarashina mini (70B)" commercialization plan within FY2025 are expected to stimulate domestic DC/GPU investment and employment, multiplying capital investment in the short term and boosting real GDP through AI capital deepening in the long term. (SB Intuitions, Inc.)
