zbeeb/Qwen2.5-Math-1.5B-Base-GRPO
zbeeb/Qwen2.5-Math-1.5B-Base-GRPO is a 1.5 billion parameter language model built upon the Qwen2.5-Math-1.5B architecture, specifically fine-tuned using the GRPO (Gradient-based Reward Policy Optimization) method. This model is the result of a token-level study focused on reinforcement learning without an initial supervised fine-tuning stage. It is primarily designed for mathematical reasoning tasks, leveraging its GRPO training on a specialized dataset to enhance problem-solving capabilities.
Loading preview...
Overview
zbeeb/Qwen2.5-Math-1.5B-Base-GRPO is a 1.5 billion parameter language model derived from the Qwen/Qwen2.5-Math-1.5B base model. This model has undergone a unique training process using Gradient-based Reward Policy Optimization (GRPO), directly from the original weights without an initial supervised fine-tuning (SFT) stage. It represents the final checkpoint from an OpenR1 token-level study focused on reinforcement learning for mathematical tasks.
Key Capabilities
- Mathematical Reasoning: Optimized for solving mathematical problems, as indicated by its base model and specialized training dataset.
- GRPO Fine-tuning: Utilizes a GRPO training recipe, which involves 1,000 optimizer updates on a dedicated mathematical dataset.
- Token-level Study: Part of a research effort to understand and improve model performance through token-level reinforcement learning.
Training Details
The model was trained using the Prime RL 0.9.0 framework on the zbeeb/Staleness-GRPO-DAPO-Math-17k dataset, comprising 17,005 rows. Key training parameters included a learning rate of 1e-06, AdamW optimizer with weight decay, and PPO ratio clipping. The training focused on 1,000 optimizer updates, with a maximum response limit of 3,072 tokens and a total context of 4,096 tokens.
Good For
- Research in Reinforcement Learning: Ideal for researchers studying GRPO, token-level optimization, and the impact of direct RL without SFT.
- Mathematical Problem Solving: Suitable for applications requiring a model with enhanced mathematical reasoning abilities, particularly where step-by-step reasoning is beneficial.
- Exploring Qwen2.5-Math Variants: Useful for developers and researchers interested in specialized versions of the Qwen2.5-Math series with unique training methodologies.