zbeeb/Qwen2.5-Math-1.5B-GRPO-Staleness-8
The zbeeb/Qwen2.5-Math-1.5B-GRPO-Staleness-8 is a 1.5 billion parameter Qwen2.5-Math model, fine-tuned using Grouped Reinforcement Learning from Policy Optimization (GRPO) with a staleness cap of 8. It is specifically optimized for mathematical reasoning tasks, trained on a 17,005-row math dataset. This model excels at solving complex math problems, providing detailed reasoning, and generating precise answers, making it suitable for applications requiring strong mathematical capabilities.
Loading preview...
Model Overview
This model, zbeeb/Qwen2.5-Math-1.5B-GRPO-Staleness-8, is a 1.5 billion parameter variant of the Qwen2.5-Math base model. It has undergone fine-tuning using Grouped Reinforcement Learning from Policy Optimization (GRPO) with a staleness cap of 8, indicating a specific training methodology to manage policy age during optimization. The model was trained on a dedicated 17,005-row math dataset, emphasizing its specialization in mathematical problem-solving.
Key Capabilities
- Mathematical Reasoning: Optimized for understanding and solving a wide range of mathematical problems.
- Detailed Explanations: Capable of generating reasoning steps leading to the final answer.
- Specific Output Formatting: Designed to produce answers in
\boxed{...}orFinal answer: ...formats. - Robust Training: Utilizes a GRPO approach with a staleness cap of 8, trained over 1,000 updates with a batch size of 64.
Performance Highlights
Evaluations on various math benchmarks demonstrate its proficiency:
- MATH500: Achieved 63.60% accuracy on
math500-pass1. - AMC23: Scored 45.00% accuracy on
amc23-pass1. - Minerva: Recorded 18.01% accuracy on
minerva-pass1.
These results reflect its performance during the training run, with a focus on mathematical equivalence for deterministic reward scoring. The model supports a total context of 4,096 tokens during training, with up to 3,072 completion tokens.
When to Use This Model
This model is particularly well-suited for applications requiring:
- Automated Math Problem Solving: Ideal for systems that need to solve and explain mathematical questions.
- Educational Tools: Can be integrated into platforms for tutoring or generating math solutions.
- Research in Mathematical AI: Useful for exploring and building upon specialized mathematical language models.