caiyuchen/DAPO-step-20
caiyuchen/DAPO-step-20 is an 8 billion parameter Qwen3-based causal language model developed by caiyuchen, fine-tuned for mathematical reasoning. This model is a training checkpoint from research on the predictability of Reinforcement Learning (RL) dynamics in LLMs, specifically demonstrating 'Rank-1 Dominance' and 'Rank-1 Linear Dynamics' in parameter updates. It is designed to support research into accelerating RL training for LLMs, achieving up to 2.5x speedup in reasoning performance. Its primary use case is for evaluating and predicting parameter dynamics during RL training, particularly for mathematical tasks.
Loading preview...
Overview
caiyuchen/DAPO-step-20 is an 8 billion parameter Qwen3-based causal language model, serving as a training checkpoint from the research paper "On Predictability of Reinforcement Learning Dynamics for Large Language Models." This model is specifically fine-tuned for mathematical reasoning tasks, leveraging the DAPO-Math-17k dataset.
Key Capabilities
- Mathematical Reasoning: Optimized for solving mathematical problems, requiring step-by-step reasoning.
- RL Dynamics Research: Provides a concrete example for studying and predicting parameter updates during Reinforcement Learning (RL) training of LLMs.
- AlphaRL Framework: Demonstrates the principles behind AlphaRL, a plug-in acceleration framework that extrapolates final parameter updates from early training, achieving up to 2.5x speedup while retaining over 96% of reasoning performance.
- Predictable Parameter Updates: Exhibits "Rank-1 Dominance" and "Rank-1 Linear Dynamics," where reasoning improvements are captured by a dominant singular subspace that evolves linearly.
Good for
- Research on LLM Training Dynamics: Ideal for researchers investigating the predictability and efficiency of RL training for large language models.
- Accelerating RL Training: Useful for understanding and implementing methods to accelerate RL-driven improvements in LLMs, particularly for reasoning tasks.
- Mathematical Problem Solving: Can be used as a base for applications requiring robust mathematical reasoning capabilities, following a specific prompt format for step-by-step reasoning.
For more details, refer to the AlphaRL GitHub repository and the associated research paper.