IamMcCoy/siwon-mini-instruct-0626
IamMcCoy's siwon-mini-instruct-0626 is a 3.8 billion parameter instruction-tuned causal language model, fine-tuned from Microsoft's Phi-4-mini-instruct. It is specifically adapted for Korean instruction-based tasks, enhancing performance through supervised fine-tuning with Korean datasets. The model features adjusted token IDs for improved instruction-following and a custom chat template for multi-turn Korean conversations. It demonstrates competitive performance on Korean language benchmarks like KMMLU and pawsx_ko.
Loading preview...
Model Overview
IamMcCoy/siwon-mini-instruct-0626 is a 3.8 billion parameter instruction-tuned model, building upon Microsoft's Phi-4-mini-instruct. Its primary focus is on enhancing performance for Korean instruction-based tasks through supervised fine-tuning with dedicated Korean datasets.
Key Differentiators & Features
- Korean Language Optimization: Specifically fine-tuned to improve instruction-following and general performance in Korean contexts.
- Token Adjustments: Addresses issues in the base model by remapping special token IDs (EOS, PAD, UNK) to unique values, preventing confusion in instruction-following tasks.
- Custom Chat Template: Incorporates an updated chat template designed to support multi-turn conversations effectively within the Korean language framework.
- Performance: Shows competitive results on Korean benchmarks such as KMMLU (0-shot) with a score of 0.3387 and pawsx_ko with 0.5485, outperforming the base Phi-4-mini-instruct in these specific metrics.
Intended Use
This model is intended for research and educational use only, with commercial use and redistribution strictly prohibited. It is particularly suitable for developers and researchers focusing on Korean natural language processing applications, especially those requiring instruction-following capabilities.