tokyotech-llm/Gemma-2-Llama-Swallow-27b-it-v0.1
The tokyotech-llm/Gemma-2-Llama-Swallow-27b-it-v0.1 is a 27 billion parameter instruction-tuned language model developed by tokyotech-llm, built upon the Gemma 2 architecture. This model was continually pre-trained on approximately 200 billion tokens, significantly enhancing its Japanese language capabilities while retaining strong English performance. It excels in multi-turn dialogue and various Japanese and English benchmarks, making it suitable for applications requiring robust bilingual understanding and generation.
Loading preview...
Overview
tokyotech-llm/Gemma-2-Llama-Swallow-27b-it-v0.1 is a 27 billion parameter instruction-tuned model from the Gemma-2-Llama-Swallow series, developed by tokyotech-llm. It is built by continually pre-training the original Gemma 2 models, focusing on enhancing Japanese language capabilities while maintaining English performance. The model was trained on approximately 200 billion tokens, including a large Japanese web corpus (Swallow Corpus Version 2), Japanese and English Wikipedia articles, and mathematical and coding content. Instruction tuning was performed using supervised fine-tuning (SFT) on synthetic data specifically designed for Japanese.
Key Capabilities
- Enhanced Japanese Language: Significantly improved performance in Japanese tasks due to extensive continual pre-training on Japanese corpora.
- Bilingual Proficiency: Retains strong English language capabilities alongside its Japanese enhancements.
- Instruction Following: Instruction-tuned for multi-turn dialogue, making it responsive and coherent in conversational contexts.
- Broad Task Performance: Demonstrates competitive performance across various Japanese benchmarks (e.g., JCommonsenseQA, JEMHopQA, JMMLU) and English benchmarks (e.g., MMLU, GSM8K, HumanEval).
Performance Highlights
On the MT-Bench JA, the 27B instruction-tuned model achieved a JMT Avg score of 0.759. For Japanese tasks, it scored 0.602 on average, with notable results in JCommonsenseQA (0.969) and JEMHopQA (0.654). In English tasks, it achieved an average score of 0.687, including 0.749 on MMLU and 0.682 on HumanEval.
When to Use This Model
This model is particularly well-suited for applications requiring high-quality, bilingual (Japanese and English) text generation and understanding. Its instruction-tuned nature makes it ideal for conversational AI, chatbots, content creation, and complex question-answering systems where strong performance in both languages is critical.