tokyotech-llm/Gemma-2-Llama-Swallow-9b-pt-v0.1
The tokyotech-llm/Gemma-2-Llama-Swallow-9b-pt-v0.1 is a 9 billion parameter pre-trained language model developed by tokyotech-llm, built upon the Gemma 2 architecture. This model enhances the Japanese language capabilities of the original Gemma 2 while retaining strong English performance, achieved through continual pre-training on approximately 200 billion tokens from Japanese web corpora, Wikipedia, and mathematical/coding content. It is designed for general language understanding and generation tasks, particularly excelling in Japanese contexts.
Loading preview...
Model Overview
The tokyotech-llm/Gemma-2-Llama-Swallow-9b-pt-v0.1 is a 9 billion parameter pre-trained language model developed by tokyotech-llm. It is part of the Gemma-2-Llama-Swallow series, which continually pre-trains on Google's Gemma 2 models to enhance Japanese language capabilities while maintaining English proficiency.
Key Capabilities & Training
- Bilingual Proficiency: Significantly improved Japanese language understanding and generation, alongside strong English performance.
- Extensive Pre-training: Continually pre-trained on approximately 200 billion tokens, including a large Japanese web corpus (Swallow Corpus Version 2), Japanese and English Wikipedia, and specialized mathematical and coding content.
- Architecture: Based on the Gemma 2 model architecture, leveraging its foundational strengths.
Performance Highlights
This model demonstrates competitive performance across various benchmarks:
- Japanese Tasks: Achieves a Ja Avg score of 0.558, outperforming
google/gemma-2-9b(0.500) andgoogle/gemma-3-12b-pt(0.518) in its size class. - English Tasks: Achieves an En Avg score of 0.595, comparable to
google/gemma-2-9b(0.597).
Good for
- Applications requiring strong bilingual (Japanese and English) language understanding and generation.
- Research and development in Japanese NLP, leveraging a robust pre-trained base.
- Tasks involving general text processing, question answering, and content generation in both languages.