laion/a3-rl-DCAgent_exp_rpt_curriculum-medium-10-8B
laion/a3-rl-DCAgent_exp_rpt_curriculum-medium-10-8B is an 8 billion parameter reinforcement learning (RL) fine-tune of the laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink model. This model was fine-tuned using SkyRL/terminus-2 on the DCAgent/exp_rpt_curriculum-medium dataset, a 512-task set, for two epochs. It is specifically optimized for tasks related to the DCAgent environment, with its checkpoint selected based on EMA-best average raw reward.
Loading preview...
Model Overview
laion/a3-rl-DCAgent_exp_rpt_curriculum-medium-10-8B is an 8 billion parameter model resulting from a reinforcement learning (RL) fine-tuning process. It is based on the laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink model, fine-tuned using the SkyRL/terminus-2 framework.
Key Training Details
- Base Model:
laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink - Fine-tuning Method: RL using SkyRL/terminus-2
- Dataset:
DCAgent/exp_rpt_curriculum-medium(a 512-task set) - Epochs: 2 epochs
- Checkpoint Selection:
global_step_10, chosen by the EMA-best ofreward/avg_raw_reward(5-step EMA, \u03b1=1/3) across the full chain. The EMA at step 10 was 0.3106 (raw reward 0.2812). - Data-limited Run: Training metrics were emitted through step 16, and checkpoints exported through step 21.
Companion Dataset and Logs
Training-time Daytona/Harbor rollouts for this run are available as a companion dataset: penfever/a3-rl-DCAgent_exp_rpt_curriculum-medium. This dataset contains the last episode of each trial, representing the rollouts the policy was trained on after rollback/truncation. Detailed per-step metrics CSVs, a metrics report, and a reward-vs-steps plot are provided in the training_logs/ directory.