mesolitica/Malaysian-Qwen2.5-1.5B-Instruct-v0.1
The mesolitica/Malaysian-Qwen2.5-1.5B-Instruct-v0.1 is a 1.5 billion parameter instruction-tuned language model, fine-tuned from Qwen/Qwen2.5-1.5B-Instruct by mesolitica. It features an extended 128k context length and is specifically optimized for understanding and generating responses in various Malaysian contexts and dialects, including Mandarin, Tamil, Jawi, and multiple regional Malay variations. This model excels in multi-turn conversations related to Malaysian legislation, politics, religions, and languages, and supports coding in these diverse linguistic contexts.
Loading preview...
Overview
mesolitica/Malaysian-Qwen2.5-1.5B-Instruct-v0.1 is a 1.5 billion parameter instruction-tuned model, building upon the Qwen/Qwen2.5-1.5B-Instruct architecture. It has been extensively fine-tuned on a highly curated 1.2 billion token Malaysian instruction dataset to enhance its understanding and generation capabilities within the Malaysian linguistic and cultural landscape.
Key Capabilities
- Extended Context Length: Features an impressive 128k context window, allowing for more comprehensive and coherent multi-turn conversations.
- Multilingual and Dialectal Support: Capable of responding and coding in a wide array of Malaysian languages and dialects, including Mandarin, Tamil, Jawi, Manglish, and various regional Malay variations (Johor, Kedah, Kelantan, Pahang, Perak, Sabah, Sarawak, Selangor, Negeri Sembilan, Terengganu).
- Malaysian Contextual Understanding: Specialized in handling multi-turn conversations related to Malaysian legislation, politics, religions, and local languages.
- Standard RAG Support: Designed to integrate effectively with Retrieval Augmented Generation (RAG) systems.
Performance
On the MalayMMLU benchmark, the model achieved an average accuracy of 54.85%, with specific category scores including 62.02% for Language and 55.61% for Humanities.
Training Details
The model was fine-tuned using LoRA with a rank of 256 and alpha of 512 (or 2.0), applying multipacking with proper SDPA causal masking to prevent document contamination. A forked CCE loss for LoRA lm_head was utilized to optimize memory consumption during training.