caiyuchen/DAPO-step-14
caiyuchen/DAPO-step-14 is an 8-billion parameter Qwen3-based causal language model developed by caiyuchen, with a 32768-token context length. This model is a training checkpoint from research on the predictability of reinforcement learning (RL) dynamics in large language models, specifically demonstrating 'Rank-1 Dominance' and 'Rank-1 Linear Dynamics' in parameter updates. It is primarily designed for research into LLM RL dynamics and mathematical reasoning tasks, particularly for evaluating and predicting how LLMs learn and improve through RL.
Loading preview...
Overview
caiyuchen/DAPO-step-14 is an 8-billion parameter Qwen3-based causal language model, serving as a training checkpoint from the research paper "On Predictability of Reinforcement Learning Dynamics for Large Language Models." This model is specifically designed to support research into the evaluation and prediction of parameter dynamics during Reinforcement Learning (RL) training of Large Language Models (LLMs). It demonstrates key properties of RL-induced parameter updates, including "Rank-1 Dominance" and "Rank-1 Linear Dynamics," which describe how reasoning improvements are captured and evolve linearly during training.
Key Capabilities
- Mathematical Reasoning: Optimized for mathematical problem-solving, as indicated by its training on the DAPO-Math-17k dataset and prompt format requiring step-by-step reasoning.
- RL Dynamics Research: Provides a concrete example for studying and predicting how LLM parameters change and improve during RL training.
- AlphaRL Framework: Contributes to the understanding and development of acceleration frameworks like AlphaRL, which extrapolates final parameter updates from early training, offering significant speedups (up to 2.5x) while maintaining high performance.
Good for
- Academic Research: Ideal for researchers investigating the internal mechanisms and predictability of RL training in LLMs.
- Mathematical Problem Solving: Suitable for applications requiring detailed, step-by-step mathematical reasoning.
- Model Analysis: Useful for analyzing parameter update dynamics and the efficiency of RL-based fine-tuning methods.