zbeeb/Qwen2.5-Math-7B-GRPO-Staleness-4
zbeeb/Qwen2.5-Math-7B-GRPO-Staleness-4 is a 7.6 billion parameter language model developed by zbeeb, fine-tuned from Qwen/Qwen2.5-Math-7B. This model is specifically optimized for mathematical reasoning tasks, utilizing a GRPO (Generalized Reinforcement Learning from Policy Optimization) training approach with a staleness cap of 4. It excels in solving complex math problems, as evidenced by its performance on benchmarks like MATH500 and AMC, making it suitable for applications requiring high-accuracy mathematical problem-solving.
Loading preview...
Overview
zbeeb/Qwen2.5-Math-7B-GRPO-Staleness-4 is a 7.6 billion parameter model derived from Qwen/Qwen2.5-Math-7B, specifically fine-tuned for mathematical problem-solving. It leverages a GRPO (Generalized Reinforcement Learning from Policy Optimization) training methodology, with a key characteristic being its 'staleness cap' of 4, which limits the age of the rollout policy during training. The model was trained on a 17,005-row DAPO math dataset, focusing on achieving high accuracy in mathematical equivalence for terminal answers.
Key Capabilities
- Advanced Mathematical Reasoning: Optimized for solving complex math problems, including those found in competitive mathematics benchmarks.
- GRPO Training: Utilizes a specialized reinforcement learning approach to enhance performance in mathematical tasks.
- Robust Evaluation: Evaluated on nine distinct mathematical datasets, including MATH500, AMC, AIME, Minerva, and OlympiadBench, demonstrating strong performance.
- Context Handling: Trained with a 4,096-token total context and supports up to 3,072 completion tokens for detailed problem explanations.
Should You Use This Model?
- For Mathematical Problem Solving: This model is ideal for applications requiring precise and accurate solutions to mathematical problems, from basic arithmetic to advanced competition-level questions.
- For Research in RLHF/GRPO: Researchers interested in the effects of GRPO training with specific staleness caps on mathematical reasoning models will find this a valuable resource.
- When High Accuracy is Critical: If your use case demands high accuracy in mathematical outputs, particularly with step-by-step reasoning, this model's specialized training makes it a strong candidate. Its performance on benchmarks like MATH500 (74.60% accuracy) and AMC (57.50% accuracy) highlights its strength in this domain.