laion/a3-rl-DCAgent_selfinstruct-naive-sandboxes-2-verified-70-8B

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 4, 2026Architecture:Transformer Featherless Exclusive Cold

The laion/a3-rl-DCAgent_selfinstruct-naive-sandboxes-2-verified-70-8B model is an 8 billion parameter language model developed by laion, fine-tuned using Reinforcement Learning (RL) with the SkyRL GRPO method. It is based on the GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink base model and trained on the DCAgent/selfinstruct-naive-sandboxes-2-verified dataset. This model is specifically designed for tasks related to agent behavior within sandbox environments, leveraging self-instructed and verified data.

Loading preview...

Model Overview

This model, a3-rl-DCAgent_selfinstruct-naive-sandboxes-2-verified-70-8B, is an 8 billion parameter language model developed by laion. It is built upon the laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink base model and has been fine-tuned using Reinforcement Learning (RL).

Key Training Details

  • Base Model: laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink
  • RL Dataset: DCAgent/selfinstruct-naive-sandboxes-2-verified, indicating a focus on self-instructed and verified data for agent training.
  • Training Method: SkyRL GRPO, utilizing 56 GPUs across 14 nodes.
  • Checkpoint: global_step_70, which represents the EMA-best checkpoint over steps 1-80 with a 5-period EMA (α=1/3).
  • Training Traces: Companion dataset penfever/a3-rl-DCAgent_selfinstruct-naive-sandboxes-2-verified contains the last episode of each trial, reflecting the rollouts the policy was trained on after rollback/truncation.

Intended Use

This model is particularly suited for research and development in areas involving agent behavior, especially within simulated or sandbox environments where self-instruction and verified outcomes are critical. Its training methodology suggests an application in tasks requiring robust and verified agent policies.