laion/a3-rl-laion_exp_rpt_stack-bash-v3-70-8B
The laion/a3-rl-laion_exp_rpt_stack-bash-v3-70-8B is an 8 billion parameter language model, fine-tuned by LAION using Reinforcement Learning (RL) on an agentic task set derived from the `exp_rpt_stack-bash-v3` dataset. This model is a partial, cancelled run based on `laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink` (Qwen3-8B) and represents an early checkpoint (global_step_70) rather than a converged model. It is preserved as a run artifact from an experimental RL series, not intended for general use due to its incomplete training and low reward signal.
Loading preview...
Overview
laion/a3-rl-laion_exp_rpt_stack-bash-v3-70-8B is an 8 billion parameter model developed by LAION, originating from an experimental Reinforcement Learning (RL) fine-tuning series. It is based on the laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink (Qwen3-8B) architecture.
Key Characteristics
- Partial Training: This model represents a partial and cancelled run from the
a3RL series, specifically checkpointglobal_step_70. - Experimental Artifact: It is preserved as a run artifact, not a fully converged or production-ready model.
- Low Reward Signal: During its limited training, the model exhibited a low and flat average raw reward (0.08-0.21), indicating non-convergence.
- Training Data: Fine-tuned on the
laion/exp_rpt_stack-bash-v3agentic task set.
Intended Use
This model is not recommended for general use cases due to its incomplete training and experimental nature. It serves primarily as a historical artifact of a cancelled research run. Researchers interested in the training traces can find companion datasets at DCAgent/a3-rl-laion_exp_rpt_stack-bash-v3.