mesolitica/Malaysian-Qwen2.5-3B-Instruct

TEXT GENERATIONConcurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Feb 23, 2025Architecture:Transformer Featherless Exclusive Cold

The mesolitica/Malaysian-Qwen2.5-3B-Instruct is a 3.1 billion parameter instruction-tuned causal language model, fine-tuned from Qwen/Qwen2.5-3B-Instruct. Developed by mesolitica, this model is specifically optimized for understanding and generating content in various Malaysian languages and dialects, including Mandarin, Tamil, Jawi, Manglish, and specific regional variations. It excels in handling multi-turn Malaysian contexts related to legislation, politics, religions, and local languages, demonstrating significant improvements over its base model on the MalayMMLU benchmark.

Loading preview...

Overview

mesolitica/Malaysian-Qwen2.5-3B-Instruct is a 3.1 billion parameter instruction-tuned model, building upon the Qwen2.5-3B-Instruct architecture. Its primary focus is on enhancing performance and understanding within the Malaysian linguistic and cultural context.

Key Capabilities

  • Multilingual and Dialectal Support: The model is fine-tuned to respond in a wide array of Malaysian languages and dialects, including Mandarin, Tamil, Jawi, Manglish, and regional variations like Johor, Kedah, Kelantan, Pahang, Perak, Sabah, Sarawak, Selangor, Negeri Sembilan, and Terengganu.
  • Contextual Understanding: It demonstrates improved comprehension of multi-turn conversations and specific Malaysian contexts, encompassing topics such as legislation, politics, religions, and local languages.
  • Code Generation: The model can generate code in the aforementioned Malaysian languages and dialects.

Performance

Benchmarking on MalayMMLU shows notable improvements over the original Qwen2.5-3B-Instruct:

  • Average Accuracy (Probability next tokens): Malaysian-Qwen2.5-3B-Instruct achieved an average accuracy of 62.21% compared to the original model's 54.59%.
  • Average Accuracy (First token match using vLLM): It scored 55.14% average accuracy, surpassing the original model's 51.06%.

Training Details

The model was fine-tuned on the mesolitica/Malaysian-SFT dataset, comprising 1.5 billion highly curated Malaysian instruction tokens. The training involved LoRA with specific configurations and a 8192 context length using multipacking and SDPA causal masking to ensure proper position IDs and prevent document contamination.