umiyuki/Llama-3-Umievo-itr014-Shizuko-8b

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:8kTool Calling:SupportedPublished:Jun 8, 2024License:llama3Architecture:Transformer0.0K Featherless Exclusive Cold

The umiyuki/Llama-3-Umievo-itr014-Shizuko-8b is an 8 billion parameter Llama-3-based model, created by umiyuki through an evolutionary merge of four Japanese-optimized Llama-3 models. This model is specifically designed for Japanese language tasks, leveraging a linear merge method to combine Meta-Llama-3-8B-Instruct, Llama-3-youko-8b-instruct-chatvector, suzume-llama-3-8B-multilingual, and shisa-v1-llama3-8b. It demonstrates strong performance in Japanese language understanding, achieving an average score of 3.85 on the ElyzaTasks100 benchmark.

Loading preview...

Model Overview

umiyuki/Llama-3-Umievo-itr014-Shizuko-8b is an 8 billion parameter language model built upon the Llama-3 architecture. It was developed by umiyuki using an evolutionary algorithm to merge four distinct Llama-3-based models, specifically chosen for their Japanese language capabilities. The merge process utilized a linear method, with meta-llama/Meta-Llama-3-8B-Instruct serving as the base model.

Key Capabilities

  • Japanese Language Proficiency: The model is explicitly designed and optimized for Japanese language understanding and generation, incorporating components from multiple Japanese-focused Llama-3 variants.
  • Merged Architecture: It combines the strengths of Meta-Llama-3-8B-Instruct, Llama-3-youko-8b-instruct-chatvector, suzume-llama-3-8B-multilingual, and shisa-v1-llama3-8b through an evolutionary merging technique.
  • Benchmark Performance: Achieved an average score of 3.85 on the ElyzaTasks100 benchmark, indicating its effectiveness in Japanese task completion.

Good For

  • Applications requiring robust Japanese language processing.
  • Developers looking for a Llama-3-based model with enhanced Japanese understanding and generation.
  • Tasks that benefit from a model specifically fine-tuned and merged for multilingual (specifically Japanese) contexts.