caiyuchen/DAPO-step-4

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Oct 3, 2025License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

caiyuchen/DAPO-step-4 is an 8 billion parameter Qwen3-based causal language model developed by Yuchen Cai et al. This model is a training checkpoint from research on predicting reinforcement learning dynamics in LLMs, specifically demonstrating Rank-1 Dominance and Rank-1 Linear Dynamics in parameter updates. It is optimized for mathematical reasoning tasks, leveraging the DAPO-Math-17k dataset and the AlphaRL acceleration framework.

Loading preview...

Overview

caiyuchen/DAPO-step-4 is an 8 billion parameter Qwen3-based causal language model, developed by Yuchen Cai et al., that serves as a key checkpoint from their research on the predictability of Reinforcement Learning (RL) dynamics in Large Language Models (LLMs). The model was trained using the DAPO-Math-17k dataset and is part of the AlphaRL framework, which aims to accelerate RL training by predicting parameter updates.

Key Characteristics

  • RL Dynamics Research: This model is specifically used to evaluate and predict parameter dynamics during RL training, demonstrating concepts like Rank-1 Dominance and Rank-1 Linear Dynamics.
  • Mathematical Reasoning: It is fine-tuned for mathematical reasoning tasks, expecting step-by-step reasoning and final answers in a boxed format.
  • AlphaRL Framework: The model is associated with AlphaRL, a plug-in acceleration framework that extrapolates final parameter updates from early training, achieving up to 2.5x speedup while retaining over 96% reasoning performance.

Intended Use

This model is primarily intended for research purposes, particularly for those studying RL dynamics in LLMs and exploring methods for accelerating RL training. It can be used to reproduce or extend the findings presented in the associated paper, "On Predictability of Reinforcement Learning Dynamics for Large Language Models" (Paper Link). The model's prompt format is designed for mathematical problem-solving, requiring specific instructions for reasoning and answer formatting.