zbeeb/Qwen2.5-Math-1.5B-Base-GRPO

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Oct 6, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

zbeeb/Qwen2.5-Math-1.5B-Base-GRPO is a 1.5 billion parameter language model built upon the Qwen2.5-Math-1.5B architecture, specifically fine-tuned using the GRPO (Gradient-based Reward Policy Optimization) method. This model is the result of a token-level study focused on reinforcement learning without an initial supervised fine-tuning stage. It is primarily designed for mathematical reasoning tasks, leveraging its GRPO training on a specialized dataset to enhance problem-solving capabilities.

Loading preview...

Overview

zbeeb/Qwen2.5-Math-1.5B-Base-GRPO is a 1.5 billion parameter language model derived from the Qwen/Qwen2.5-Math-1.5B base model. This model has undergone a unique training process using Gradient-based Reward Policy Optimization (GRPO), directly from the original weights without an initial supervised fine-tuning (SFT) stage. It represents the final checkpoint from an OpenR1 token-level study focused on reinforcement learning for mathematical tasks.

Key Capabilities

  • Mathematical Reasoning: Optimized for solving mathematical problems, as indicated by its base model and specialized training dataset.
  • GRPO Fine-tuning: Utilizes a GRPO training recipe, which involves 1,000 optimizer updates on a dedicated mathematical dataset.
  • Token-level Study: Part of a research effort to understand and improve model performance through token-level reinforcement learning.

Training Details

The model was trained using the Prime RL 0.9.0 framework on the zbeeb/Staleness-GRPO-DAPO-Math-17k dataset, comprising 17,005 rows. Key training parameters included a learning rate of 1e-06, AdamW optimizer with weight decay, and PPO ratio clipping. The training focused on 1,000 optimizer updates, with a maximum response limit of 3,072 tokens and a total context of 4,096 tokens.

Good For

  • Research in Reinforcement Learning: Ideal for researchers studying GRPO, token-level optimization, and the impact of direct RL without SFT.
  • Mathematical Problem Solving: Suitable for applications requiring a model with enhanced mathematical reasoning abilities, particularly where step-by-step reasoning is beneficial.
  • Exploring Qwen2.5-Math Variants: Useful for developers and researchers interested in specialized versions of the Qwen2.5-Math series with unique training methodologies.