zbeeb/Qwen2.5-3B-GRPO-Staleness-8

TEXT GENERATIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:3.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 19, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

zbeeb/Qwen2.5-3B-GRPO-Staleness-8 is a 3.1 billion parameter causal language model based on the Qwen2.5 architecture, fine-tuned using the GRPO method. This model is specifically optimized for mathematical reasoning tasks, trained on a 17,005-row DAPO math dataset. It features a 32,768-token context length and demonstrates capabilities in solving complex math problems with detailed reasoning. Its primary strength lies in its specialized mathematical problem-solving abilities, making it suitable for applications requiring precise numerical and logical computations.

Loading preview...

Model Overview

zbeeb/Qwen2.5-3B-GRPO-Staleness-8 is a 3.1 billion parameter model derived from the Qwen/Qwen2.5-3B base, specifically fine-tuned using the GRPO (Generalized Reinforcement Learning with Policy Optimization) method. This model was trained on a dedicated 17,005-row DAPO math dataset, emphasizing mathematical problem-solving and reasoning. A key aspect of its training involved a 'staleness cap' of 8 (max_off_policy_steps), which limits the age of the rollout policy during the 1,000 training updates.

Key Capabilities

  • Mathematical Reasoning: Optimized for solving complex math problems, as evidenced by its training on a specialized math dataset and evaluation on benchmarks like MATH500, AMC, and AIME.
  • Detailed Explanations: Designed to provide reasoning alongside answers, with a prompt format encouraging explanations before a final boxed or line answer.
  • Context Length: Supports a substantial context window of 32,768 tokens, allowing for processing longer problem descriptions or multi-step reasoning.
  • Robust Training: Underwent 1,000 updates with a batch size of 64 and a deterministic reward system for mathematical equivalence, ensuring focused optimization.

Performance Highlights

Evaluations on various math benchmarks show its performance in mathematical problem-solving. For instance, it achieved 59.40% accuracy on math500-pass1 and 32.50% on amc23-pass1. These scores reflect its specialized training for mathematical tasks.

When to Use This Model

This model is particularly well-suited for applications requiring:

  • Automated Math Solvers: Ideal for systems that need to compute and explain solutions to mathematical problems.
  • Educational Tools: Can be integrated into platforms for generating step-by-step solutions or checking student work.
  • Research in Mathematical AI: Useful for exploring the capabilities of LLMs in complex numerical and logical reasoning.

It's important to note that while the model has a 32K context length, the training used a 4,096-token total context for completions, and 8K context extension was not configured or validated in this release.