laion/a3-rl-DCAgent_selfinstruct-naive-sandboxes-2-verified-70-8B
The laion/a3-rl-DCAgent_selfinstruct-naive-sandboxes-2-verified-70-8B model is an 8 billion parameter language model developed by laion, fine-tuned using Reinforcement Learning (RL) with the SkyRL GRPO method. It is based on the GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink base model and trained on the DCAgent/selfinstruct-naive-sandboxes-2-verified dataset. This model is specifically designed for tasks related to agent behavior within sandbox environments, leveraging self-instructed and verified data.
Loading preview...
Model Overview
This model, a3-rl-DCAgent_selfinstruct-naive-sandboxes-2-verified-70-8B, is an 8 billion parameter language model developed by laion. It is built upon the laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink base model and has been fine-tuned using Reinforcement Learning (RL).
Key Training Details
- Base Model:
laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink - RL Dataset:
DCAgent/selfinstruct-naive-sandboxes-2-verified, indicating a focus on self-instructed and verified data for agent training. - Training Method: SkyRL GRPO, utilizing 56 GPUs across 14 nodes.
- Checkpoint:
global_step_70, which represents the EMA-best checkpoint over steps 1-80 with a 5-period EMA (α=1/3). - Training Traces: Companion dataset
penfever/a3-rl-DCAgent_selfinstruct-naive-sandboxes-2-verifiedcontains the last episode of each trial, reflecting the rollouts the policy was trained on after rollback/truncation.
Intended Use
This model is particularly suited for research and development in areas involving agent behavior, especially within simulated or sandbox environments where self-instruction and verified outcomes are critical. Its training methodology suggests an application in tasks requiring robust and verified agent policies.