laion/a3-rl-DCAgent_exp_rpt_e2egit-large-15-8B
The laion/a3-rl-DCAgent_exp_rpt_e2egit-large-15-8B is an 8 billion parameter Reinforcement Learning (RL) checkpoint, specifically trained using GRPO/rloo_n via SkyRL. It was fine-tuned from a GLM-4 base model on the DCAgent/exp_rpt_e2egit-large dataset. This model represents global_step_15, identified as the optimal checkpoint based on its reward metrics before a late-stage training collapse, making it suitable for tasks requiring robust performance from an early-mid training phase.
Loading preview...
Model Overview
The laion/a3-rl-DCAgent_exp_rpt_e2egit-large-15-8B is an 8 billion parameter Reinforcement Learning (RL) model checkpoint. It was developed by laion and trained using the GRPO/rloo_n algorithm via SkyRL, building upon the laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink base model. The training utilized the DCAgent/exp_rpt_e2egit-large dataset.
Key Characteristics
- Optimal Checkpoint Selection: This model represents
global_step_15, which was identified as the best performing checkpoint. This selection was based on the 5-period Exponential Moving Average (EMA) ofreward/avg_raw_reward(EMA 0.864, raw 0.842), chosen because the full 80-step training run experienced a performance collapse in later stages (entropy decreased from 0.0961 to 0.0215 after approximately step 56). - Training Configuration: The model was trained using the
hpc/skyrl_yaml/jupiter/56GPU_base.yamllaunch configuration, which is included asrl_config.yaml.
Training Details
- Training Traces: Training-time Daytona/Harbor rollouts are currently deferred due to I/O issues during the
make_and_upload_trace_datasetscript execution. The companion dataset is planned to be available at open-athena/a3-rl-DCAgent_exp_rpt_e2egit-large once re-attempted. - Training Metrics: Detailed training logs, including per-chain-link
metrics_job_*.csv, mergedmetrics_table.csvandvllm_metrics_table.csv, a markdown report (metrics_report.md), and reward/turn-count plots, are provided in thetraining_logs/directory. Raw.outlogs from all 10 chain links are also included.