tokyotech-llm/Gemma-2-Llama-Swallow-9b-it-v0.1
The tokyotech-llm/Gemma-2-Llama-Swallow-9b-it-v0.1 is a 9 billion parameter instruction-tuned language model developed by tokyotech-llm, built upon the Gemma 2 architecture. This model enhances Japanese language capabilities through continual pre-training on approximately 200 billion tokens from Japanese web corpora, while retaining strong English language performance. It is specifically fine-tuned using synthetic Japanese data, making it particularly strong for Japanese multi-turn dialogue and various Japanese NLP tasks.
Loading preview...
Overview
The tokyotech-llm/Gemma-2-Llama-Swallow-9b-it-v0.1 is a 9 billion parameter instruction-tuned model developed by tokyotech-llm. It is part of the Gemma-2-Llama-Swallow series, which is built by continually pre-training on Google's Gemma 2 models. The primary focus of this series is to significantly enhance Japanese language capabilities while maintaining strong performance in English.
Key Capabilities
- Enhanced Japanese Language Processing: The model underwent continual pre-training on approximately 200 billion tokens sampled from a large Japanese web corpus (Swallow Corpus Version 2), Japanese and English Wikipedia articles, and mathematical/coding content.
- Instruction-Tuned for Japanese: This specific model variant (
-it) was fine-tuned using supervised learning on synthetic data specially created for Japanese instructions, making it adept at understanding and generating Japanese responses. - Bilingual Performance: While optimized for Japanese, it retains the English language capabilities of its base Gemma 2 architecture.
- Strong Benchmark Performance: On the Japanese MT-Bench, the 9B instruction-tuned model achieves a JMT Avg score of 0.749, outperforming
google/gemma-2-9b-it(0.736). It also shows competitive performance across various Japanese and English evaluation benchmarks, including JCommonsenseQA, JSQuAD, MMLU, and GSM8K.
Use Cases
This model is particularly well-suited for applications requiring robust performance in both Japanese and English, especially for:
- Japanese Dialogue Systems: Its instruction-tuning on synthetic Japanese data makes it effective for multi-turn conversations in Japanese.
- Multilingual NLP Tasks: Ideal for tasks involving Japanese and English text generation, summarization, and question answering.
- Research and Development: Provides a strong base for further fine-tuning on specific Japanese-centric applications.