caiyuchen/DAPO-step-1

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Oct 3, 2025License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The caiyuchen/DAPO-step-1 is an 8 billion parameter Qwen3-based language model developed by caiyuchen, fine-tuned for mathematical reasoning tasks. This model is a training checkpoint from research exploring the predictability of reinforcement learning dynamics in LLMs, specifically demonstrating 'Rank-1 Dominance' and 'Rank-1 Linear Dynamics' in parameter updates. It is designed to support research into accelerating RL training for LLMs, offering insights into how reasoning improvements manifest during fine-tuning. The model is part of the AlphaRL framework, which aims to achieve significant speedups in RL training while maintaining performance.

Loading preview...

Model Overview

The caiyuchen/DAPO-step-1 is an 8 billion parameter language model built upon the Qwen3-8B-Base architecture. Developed by caiyuchen, this model serves as a specific training checkpoint from the research paper "On Predictability of Reinforcement Learning Dynamics for Large Language Models." Its primary purpose is to facilitate the study and evaluation of reinforcement learning (RL) dynamics within large language models (LLMs), particularly concerning mathematical reasoning tasks.

Key Research Insights & Capabilities

This model embodies the findings of the associated research, which identifies two crucial properties of RL-induced parameter updates:

  • Rank-1 Dominance: The top singular subspace of the parameter update matrix captures nearly all improvements in reasoning capabilities.
  • Rank-1 Linear Dynamics: This dominant subspace evolves linearly throughout the training process, enabling accurate prediction from early checkpoints.

These insights led to the development of AlphaRL, a plug-in acceleration framework. AlphaRL extrapolates final parameter updates from a short early training window, achieving up to 2.5x speedup while retaining over 96% of reasoning performance. This model is a component used in validating these findings.

Use Cases

  • Research on RL Dynamics: Ideal for researchers studying parameter dynamics, predictability, and acceleration techniques in RL training for LLMs.
  • Mathematical Reasoning Evaluation: Can be used to evaluate and benchmark mathematical reasoning capabilities, particularly when comparing different stages of RL fine-tuning.
  • AlphaRL Framework Exploration: Provides a practical example for understanding and implementing the AlphaRL acceleration framework.