Jinhe/ReflectRL-Qwen2.5-Math-7B-DAPO

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 4, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Jinhe/ReflectRL-Qwen2.5-Math-7B-DAPO is a Qwen2.5-based model developed by Jinhe Bi and collaborators, fine-tuned using the ReflectRL framework. This model specializes in mathematical reasoning tasks by learning from Golden Negative Trajectories (GNTs) through a reflective-to-direct reasoning approach. It leverages on-policy post-training to improve performance in problem-solving by initially using GNTs as reflective context and then transitioning to direct inference.

Loading preview...

ReflectRL-Qwen2.5-Math-7B-DAPO Overview

This model, developed by Jinhe Bi and collaborators, is a specialized Qwen2.5-based checkpoint that integrates the ReflectRL framework for enhanced mathematical reasoning. ReflectRL is a novel, lightweight framework designed for on-policy post-training, which uniquely learns from Golden Negative Trajectories (GNTs).

Key Capabilities & Training Approach

  • Reflective-to-Direct Reasoning: ReflectRL utilizes GNTs not for direct imitation, but as a reflective context during training. This allows the model to understand and learn from failed expert trajectories.
  • Gradual Policy Transition: The training process gradually shifts the policy from relying on reflective context to performing direct reasoning during inference, optimizing for efficient problem-solving.
  • Mathematical Task Specialization: The model is specifically fine-tuned to excel in mathematical reasoning, leveraging its unique training methodology to improve accuracy and robustness in this domain.

Good For

  • Researchers and developers interested in advanced reasoning techniques, particularly those involving learning from negative examples.
  • Applications requiring robust mathematical problem-solving capabilities.
  • Exploring novel on-policy post-training methods for language models.