ExploreXploitQ/DeepMath-103K-Level6-Qwen3-4B-Base-GRPO
ExploreXploitQ/DeepMath-103K-Level6-Qwen3-4B-Base-GRPO is a 4 billion parameter language model based on Qwen3-4B-Base, specifically post-trained for advanced mathematical reasoning. It utilizes Group Relative Policy Optimization (GRPO) on the Level 6 subset of the DeepMath-103K dataset. This model is optimized for solving complex math problems, demonstrated by its performance on AIME and AMC validation sets. With a context length of 32,768 tokens, it is designed for applications requiring deep mathematical problem-solving capabilities.
Loading preview...
Model Overview
ExploreXploitQ/DeepMath-103K-Level6-Qwen3-4B-Base-GRPO is a specialized 4-billion parameter language model derived from Qwen/Qwen3-4B-Base. Its primary distinction lies in its post-training methodology: it was fine-tuned using Group Relative Policy Optimization (GRPO) on the challenging Level 6 subset of the DeepMath-103K dataset.
Key Capabilities and Training
This model is engineered for advanced mathematical reasoning. The training involved 57,046 examples from DeepMath Level 6, utilizing the verl framework and full-parameter actor updates with rule-based math outcome rewards. Key training parameters include a context length of 32,768 tokens, 8 responses per prompt during rollout, and a constant learning rate of 1e-6 over one epoch.
Performance Highlights
Validation was conducted on competitive math datasets, yielding the following average accuracies at 16 responses:
- AIME 2025: 21.25%
- AMC 2022–2023: 60.77%
- AIME 2024: 23.54%
Use Cases
This model is particularly well-suited for applications requiring robust mathematical problem-solving, especially those involving complex, multi-step reasoning found in advanced mathematics competitions or educational tools. Its specialized training makes it a strong candidate for tasks where precise mathematical output and step-by-step reasoning are critical.