laion/ablation-pymethods2test-seqmean-arm0-15-8B
The laion/ablation-pymethods2test-seqmean-arm0-15-8B is an 8 billion parameter language model, based on a Qwen3-8B SFT model, developed by laion. This model is a Reinforcement Learning (RL) checkpoint from an ablation study focusing on sequence-mean policy loss with the RLOO-n advantage estimator, contrasting with token-mean reduction. It was trained using SkyRL GRPO on the DCAgent/exp_rpt_pymethods2test-large dataset, making it suitable for research into RL training methodologies and their impact on model performance.
Loading preview...
Model Overview
laion/ablation-pymethods2test-seqmean-arm0-15-8B is an 8 billion parameter language model derived from a Qwen3-8B SFT base model. It represents a specific checkpoint (global_step_15) from a Reinforcement Learning (RL) ablation study conducted by laion, focusing on the impact of different policy loss reduction methods.
Key Characteristics
- RL Training: This model was trained using the SkyRL GRPO (Generalized Policy Regularization Optimization) algorithm.
- Ablation Study Focus: It is part of a study investigating the
sequence-meanpolicy loss reduction with theRLOO-n(RLOO-n) advantage estimator, a departure from thetoken-meanreduction used in other series. - Base Model: The foundation for this model is
laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink, which is a Qwen3-8B SFT variant. - Training Data: The model was trained on the
DCAgent/exp_rpt_pymethods2test-largedataset. - Reproducibility: The exact
rl_config.jsonused for its training is included in the repository for full reproducibility.
Training Details
The training involved 14x GH200 nodes on JSC Jupiter. Companion training traces, specifically Daytona/Harbor rollouts, are available as a separate dataset: open-athena/ablation-pymethods2test-seqmean-arm0. These traces contain the last episode of each trial, reflecting the rollouts the policy was trained on after rollback or truncation.
Use Cases
This model is particularly valuable for researchers and developers interested in:
- Exploring the effects of different RL policy loss reduction techniques (e.g., sequence-mean vs. token-mean).
- Analyzing the performance and characteristics of models trained with the RLOO-n advantage estimator.
- Reproducing and extending RL ablation studies in large language models.
- Understanding the training dynamics and trace data of RL-tuned models.