umiyuki/Llama-3-Umievo-itr014-Shizuko-8b
The umiyuki/Llama-3-Umievo-itr014-Shizuko-8b is an 8 billion parameter Llama-3-based model, created by umiyuki through an evolutionary merge of four Japanese-optimized Llama-3 models. This model is specifically designed for Japanese language tasks, leveraging a linear merge method to combine Meta-Llama-3-8B-Instruct, Llama-3-youko-8b-instruct-chatvector, suzume-llama-3-8B-multilingual, and shisa-v1-llama3-8b. It demonstrates strong performance in Japanese language understanding, achieving an average score of 3.85 on the ElyzaTasks100 benchmark.
Loading preview...
Model Overview
umiyuki/Llama-3-Umievo-itr014-Shizuko-8b is an 8 billion parameter language model built upon the Llama-3 architecture. It was developed by umiyuki using an evolutionary algorithm to merge four distinct Llama-3-based models, specifically chosen for their Japanese language capabilities. The merge process utilized a linear method, with meta-llama/Meta-Llama-3-8B-Instruct serving as the base model.
Key Capabilities
- Japanese Language Proficiency: The model is explicitly designed and optimized for Japanese language understanding and generation, incorporating components from multiple Japanese-focused Llama-3 variants.
- Merged Architecture: It combines the strengths of
Meta-Llama-3-8B-Instruct,Llama-3-youko-8b-instruct-chatvector,suzume-llama-3-8B-multilingual, andshisa-v1-llama3-8bthrough an evolutionary merging technique. - Benchmark Performance: Achieved an average score of 3.85 on the ElyzaTasks100 benchmark, indicating its effectiveness in Japanese task completion.
Good For
- Applications requiring robust Japanese language processing.
- Developers looking for a Llama-3-based model with enhanced Japanese understanding and generation.
- Tasks that benefit from a model specifically fine-tuned and merged for multilingual (specifically Japanese) contexts.