pb09204048/Qwen3-4B-Math-RL-Specialist
The pb09204048/Qwen3-4B-Math-RL-Specialist is a 4 billion parameter BF16 math specialist model, fine-tuned from Qwen/Qwen3-4B using a rank-16 LoRA. It is specifically optimized for mathematical reasoning tasks, demonstrating a significant performance increase on the AIME24 benchmark. This model is designed to serve as a domain teacher for Multi-Teacher On-Policy Distillation (MOPD) experiments, focusing on mathematical problem-solving.
Loading preview...
Overview
This model, pb09204048/Qwen3-4B-Math-RL-Specialist, is a 4 billion parameter BF16 math specialist derived from the Qwen3-4B base model. It was fine-tuned using a rank-16 LoRA and merged into full Hugging Face weights. The primary purpose of this specialist is to act as a domain teacher for Multi-Teacher On-Policy Distillation (MOPD) experiments, specifically in the mathematical domain.
Key Capabilities & Training
- Mathematical Specialization: Trained extensively on the
zhuzilin/dapo-math-17kdataset, comprising 15,981 unique math prompts. - Performance: Achieved a 50.00% pass@1 score on the AIME24 benchmark, a substantial improvement over the Qwen3-4B base model's 22.50%.
- Training Method: Utilizes independent domain Reinforcement Learning (RL) with GRPO-style group-centered advantages and a PPO clipped policy objective. The training involved 500 optimizer updates.
- Context Length: Supports a maximum response length of 16,384 tokens and a 24,576-token context during evaluation.
Use Cases
This model is particularly suited for:
- MOPD Experiments: Designed to be used as a math teacher model within a MOPD framework, providing token-level supervision for student models.
- Mathematical Reasoning: Excels at solving complex mathematical problems, as evidenced by its AIME24 performance.
- Research: Ideal for researchers exploring domain-specific RL fine-tuning and multi-teacher distillation techniques.