laion/a3-rl-DCAgent_exp_rpt_curriculum-medium-10-8B

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 5, 2026Architecture:Transformer Featherless Exclusive Cold

laion/a3-rl-DCAgent_exp_rpt_curriculum-medium-10-8B is an 8 billion parameter reinforcement learning (RL) fine-tune of the laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink model. This model was fine-tuned using SkyRL/terminus-2 on the DCAgent/exp_rpt_curriculum-medium dataset, a 512-task set, for two epochs. It is specifically optimized for tasks related to the DCAgent environment, with its checkpoint selected based on EMA-best average raw reward.

Loading preview...

Model Overview

laion/a3-rl-DCAgent_exp_rpt_curriculum-medium-10-8B is an 8 billion parameter model resulting from a reinforcement learning (RL) fine-tuning process. It is based on the laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink model, fine-tuned using the SkyRL/terminus-2 framework.

Key Training Details

  • Base Model: laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink
  • Fine-tuning Method: RL using SkyRL/terminus-2
  • Dataset: DCAgent/exp_rpt_curriculum-medium (a 512-task set)
  • Epochs: 2 epochs
  • Checkpoint Selection: global_step_10, chosen by the EMA-best of reward/avg_raw_reward (5-step EMA, \u03b1=1/3) across the full chain. The EMA at step 10 was 0.3106 (raw reward 0.2812).
  • Data-limited Run: Training metrics were emitted through step 16, and checkpoints exported through step 21.

Companion Dataset and Logs

Training-time Daytona/Harbor rollouts for this run are available as a companion dataset: penfever/a3-rl-DCAgent_exp_rpt_curriculum-medium. This dataset contains the last episode of each trial, representing the rollouts the policy was trained on after rollback/truncation. Detailed per-step metrics CSVs, a metrics report, and a reward-vs-steps plot are provided in the training_logs/ directory.