jaeyong2/Qwen2.5-0.5B-Instruct-Ja-SFT
The jaeyong2/Qwen2.5-0.5B-Instruct-Ja-SFT is a 0.5 billion parameter instruction-tuned language model, based on the Qwen2.5 architecture, specifically fine-tuned for Japanese language tasks. It demonstrates competitive performance on Japanese evaluation benchmarks like llm-jp-eval, particularly in tasks such as reading comprehension (RC) and machine translation (MT). This model is optimized for efficient deployment in Japanese natural language processing applications requiring a compact yet capable instruction-following model.
Loading preview...
Model Overview
The jaeyong2/Qwen2.5-0.5B-Instruct-Ja-SFT is a compact 0.5 billion parameter instruction-tuned model, building upon the Qwen2.5 architecture. This specific variant has been fine-tuned for Japanese language understanding and generation, making it suitable for various Japanese NLP tasks.
Key Capabilities & Performance
The model's performance is evaluated using the llm-jp-eval script, which assesses its capabilities across multiple Japanese NLP benchmarks. Key performance indicators include:
- Reading Comprehension (RC): Achieves 0.5847 on the
llm-jp-evalRC task after fine-tuning. - Machine Translation (MT): Shows strong performance with a score of 0.6691 on the
llm-jp-evalMT task. - MMLU: Scores 0.4614, indicating general knowledge and reasoning abilities.
While some areas like Common Sense Generation (CG) and Factuality (FA) show lower scores, the model demonstrates notable improvements in areas like Elementary Logic (EL) and Multiple Choice (MC) after fine-tuning.
License
This model is based on the Qwen/Qwen2.5-0.5B-Instruct, which is licensed under the Apache-2.0 License.
Acknowledgement
Development of this model was supported by the TPU Research Cloud program.