tokyotech-llm/Llama-3.1-Swallow-8B-Instruct-v0.5
Llama-3.1-Swallow-8B-Instruct-v0.5 is an 8 billion parameter instruction-tuned causal language model developed by tokyotech-llm. Built upon Meta Llama 3.1, it significantly enhances Japanese language capabilities through continual pre-training on a 200 billion token Japanese web corpus while retaining strong English performance. This model excels in Japanese multi-turn dialogue, achieving state-of-the-art performance on Japanese MT-Bench among open-source LLMs of comparable size.
Loading preview...
Llama 3.1 Swallow 8B Instruct v0.5: Enhanced Japanese LLM
This model is an 8 billion parameter instruction-tuned variant of the Llama 3.1 Swallow series, developed by tokyotech-llm. It is built by continually pre-training the Meta Llama 3.1 base model, specifically enhancing its Japanese language capabilities while maintaining strong English performance.
Key Capabilities & Differentiators
- Bilingual Proficiency: Significantly improved Japanese language understanding and generation, alongside robust English capabilities.
- Continual Pre-training: Utilizes approximately 200 billion tokens from a large Japanese web corpus (Swallow Corpus Version 2), Japanese and English Wikipedia, and mathematical/coding content.
- Instruction Tuning: Fine-tuned with synthetic data specifically designed for Japanese instructions, imitating the conversational behavior of gemma-3-27b-it.
- Leading Japanese MT-Bench Performance: Achieves state-of-the-art results on Japanese MT-Bench among open-source LLMs with up to 8 billion parameters, outperforming its predecessors.
Ideal Use Cases
- Applications requiring high-quality Japanese language generation and understanding.
- Multi-turn dialogue systems in Japanese.
- Tasks benefiting from a strong bilingual foundation in both Japanese and English.
Top 3 parameter combinations used by Featherless users for this model. Click a tab to see each config.