Pradheep1647/q0-gsm8k-1_5b-mopd
Pradheep1647/q0-gsm8k-1_5b-mopd is a 1.5 billion parameter language model based on Qwen2.5-1.5B-Instruct, fine-tuned using Multi-Teacher On-Policy Distillation (MOPD). This model is specifically optimized for mathematical reasoning tasks, demonstrating strong performance on the GSM8K benchmark with a pass@8 of 84.77%. It is designed for efficient deployment, with its LoRA adapter merged into the 16-bit base weights.
Loading preview...
Model Overview
Pradheep1647/q0-gsm8k-1_5b-mopd is a 1.5 billion parameter model derived from Qwen/Qwen2.5-1.5B-Instruct. Its primary distinction lies in its fine-tuning methodology: Multi-Teacher On-Policy Distillation (MOPD). This technique involved distilling knowledge from a q0 cyclic-trajectory GRPO mixture, utilizing two top-performing q0 snapshots (cycle02 + cycle03) as teachers.
Key Capabilities and Performance
This model is specifically engineered for mathematical reasoning. Its performance on relevant benchmarks highlights this specialization:
- GSM8K Validation (256-example slice):
- Pass@1: 49.07%
- Pass@4: 76.29%
- Pass@8: 84.77%
- MATH-500 Test (500 problems):
- Pass@1: 16.40%
- Pass@4: 35.00%
Technical Details
The model was trained using 4-bit QLoRA (NF4, r=8, alpha=16) and the LoRA adapter has been merged into the 16-bit base weights, making it directly deployable. This training was conducted on an RTX 4060 Laptop (8 GB).