rinna/llama-3-youko-70b
The rinna/llama-3-youko-70b is a 70 billion parameter Llama 3-based language model developed by rinna, continually pre-trained on an additional 5 billion tokens of mixed Japanese and English datasets. This model significantly enhances performance on Japanese language tasks while retaining the original Llama 3 architecture and tokenizer. It is optimized for applications requiring strong Japanese language understanding and generation capabilities.
Loading preview...
Overview
rinna/llama-3-youko-70b is a 70 billion parameter language model built upon Meta's Llama 3-70B. Developed by rinna, this model underwent continual pre-training on approximately 5 billion additional tokens comprising a mixture of Japanese and English datasets. This targeted pre-training significantly boosts its proficiency in Japanese language tasks.
Key Capabilities
- Enhanced Japanese Performance: Achieves improved performance on Japanese benchmarks due to extensive continual pre-training on diverse Japanese corpora, including Japanese CC-100, C4, OSCAR, and rinna's curated datasets.
- Llama 3 Foundation: Retains the robust architecture and tokenization of the original Meta Llama 3-70B model, ensuring strong general language understanding.
- Large Scale: With 70 billion parameters, it offers substantial capacity for complex language tasks.
Good For
- Japanese Language Applications: Ideal for use cases requiring high-quality Japanese text generation, comprehension, and analysis.
- Research and Development: Suitable for researchers and developers exploring multilingual LLMs, particularly those focusing on Japanese language processing.
- Building upon Llama 3: Provides a strong foundation for further fine-tuning or instruction-tuning for specific Japanese-centric applications.