caiyuchen/DAPO-step-3

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Oct 3, 2025License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The caiyuchen/DAPO-step-3 model is an 8 billion parameter Qwen3-based language model developed by caiyuchen, fine-tuned for mathematical reasoning. This model is a specific training checkpoint used in research on predicting reinforcement learning dynamics in LLMs. It is designed to evaluate and predict how parameter updates during RL training impact reasoning improvements, particularly in mathematical tasks. Its primary differentiator is its role in demonstrating 'Rank-1 Dominance' and 'Rank-1 Linear Dynamics' in RL-induced parameter updates.

Loading preview...

Overview

This model, caiyuchen/DAPO-step-3, is an 8 billion parameter Qwen3-based language model developed by caiyuchen. It represents a specific training checkpoint from the research paper "On Predictability of Reinforcement Learning Dynamics for Large Language Models." The core focus of this research is to understand and predict how Reinforcement Learning (RL) influences the parameter dynamics of Large Language Models (LLMs), especially concerning reasoning capabilities.

Key Insights & Capabilities

The research identifies two crucial properties of RL-induced parameter updates:

  • Rank-1 Dominance: The most significant improvements in reasoning are captured by the top singular subspace of the parameter update matrix.
  • Rank-1 Linear Dynamics: This dominant subspace evolves linearly throughout training, enabling accurate predictions from early training stages.

These insights led to the development of AlphaRL, an acceleration framework that can extrapolate final parameter updates from a short initial training period. This framework achieves up to a 2.5x speedup while retaining over 96% of the original reasoning performance. This model is provided to facilitate further research into these RL dynamics.

Prompt Format

For inference, questions should be formatted with the instruction "Please reason step by step, and put your final answer within boxed{}" and then wrapped using the standard chat template for Qwen3 models.

Good for

  • Researchers studying the dynamics of Reinforcement Learning in LLMs.
  • Evaluating methods for predicting parameter updates during RL training.
  • Understanding the principles behind the AlphaRL acceleration framework for LLM training.