laion/a3-rl-DCAgent_exp_rpt_pymethods2test-v3-10-8B
The laion/a3-rl-DCAgent_exp_rpt_pymethods2test-v3-10-8B is an 8 billion parameter reinforcement learning (RL) checkpoint, specifically a SkyRL model. It was trained from the laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink base model using the DCAgent/exp_rpt_pymethods2test-v3 dataset. This model is optimized for tasks related to the DCAgent environment, with its checkpoint selected at global_step 10 based on a 5-period reward-EMA.
Loading preview...
Model Overview
This model, a3-rl-DCAgent_exp_rpt_pymethods2test-v3-10-8B, is an 8 billion parameter reinforcement learning (RL) checkpoint developed by laion. It utilizes the SkyRL framework and was fine-tuned from the base model laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink.
Training Details
The model was trained on the DCAgent/exp_rpt_pymethods2test-v3 dataset. The specific checkpoint provided was selected at global_step 10, based on a 5-period reward-EMA (exponential moving average with alpha=1/3) calculated over the full stitched step sequence. This selection was made to avoid degenerate resume-creep links observed in later steps (15-25), ensuring the chosen checkpoint represents genuinely trained and aligned performance.
Companion Data
Training-time Daytona/Harbor rollouts associated with this run are available as a companion dataset: open-athena/a3-rl-DCAgent_exp_rpt_pymethods2test-v3. This dataset contains the last episode of each trial, representing the same rollouts the policy was trained on after rollback and truncation.
Use Cases
This model is primarily intended for research and development within the DCAgent environment, particularly for tasks related to the exp_rpt_pymethods2test-v3 dataset. It can be used to analyze RL policy behavior and performance within this specific domain.