tokyotech-llm/Llama-3.1-Swallow-8B-v0.5
Llama 3.1 Swallow 8B v0.5 is an 8 billion parameter large language model developed by tokyotech-llm, built by continually pre-training Meta Llama 3.1. This model significantly enhances Japanese language capabilities, reasoning in code and math, while maintaining strong English performance. It was trained on approximately 210 billion tokens from a diverse corpus including Japanese web data, Wikipedia, and specialized mathematical and coding content, making it particularly strong for bilingual applications requiring robust reasoning.
Loading preview...
Llama 3.1 Swallow 8B v0.5: Enhanced Japanese and Reasoning
Llama 3.1 Swallow 8B v0.5 is an 8 billion parameter large language model developed by tokyotech-llm, based on continual pre-training of Meta's Llama 3.1. This version specifically targets and achieves significant improvements in Japanese language capabilities and reasoning tasks (code and math), while preserving the strong English language performance of the original Llama 3.1.
Key Enhancements & Training
The model underwent continual pre-training on approximately 210 billion tokens. This extensive dataset includes:
- Swallow Corpus Version 2: A large, high-quality Japanese web corpus.
- Japanese and English Wikipedia articles.
- Specialized content for code and mathematics: Including proprietary datasets like Swallow Code Version 1 and Swallow Math Version 1, which have been shown to outperform other datasets in boosting LLM performance in these domains.
Instruction-tuned variants (Instruct models) were further refined using synthetic data tailored for Japanese.
Performance Highlights
Benchmarking demonstrates that Llama 3.1 Swallow 8B v0.5 achieves competitive and often superior results compared to other 8B models in Japanese tasks. For instance, it shows strong performance in JCommonsenseQA, NIILC, and JSQuAD, and improved scores in MGSM and JHumanEval over the base Llama 3.1 8B. English task performance is largely maintained, with notable improvements in GSM8K and MATH compared to the base Llama 3.1 8B.
Ideal Use Cases
This model is particularly well-suited for applications requiring:
- High-quality Japanese language generation and understanding.
- Bilingual (Japanese-English) applications where strong performance in both languages is critical.
- Tasks involving mathematical and coding reasoning.
- Developers seeking a Llama 3.1-based model with enhanced specialized capabilities for the Japanese market and technical domains.