laion/a3-rl-laion_exp_rpt_stack-bash-v3-70-8B

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 6, 2026Architecture:Transformer Featherless Exclusive Cold

The laion/a3-rl-laion_exp_rpt_stack-bash-v3-70-8B is an 8 billion parameter language model, fine-tuned by LAION using Reinforcement Learning (RL) on an agentic task set derived from the `exp_rpt_stack-bash-v3` dataset. This model is a partial, cancelled run based on `laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink` (Qwen3-8B) and represents an early checkpoint (global_step_70) rather than a converged model. It is preserved as a run artifact from an experimental RL series, not intended for general use due to its incomplete training and low reward signal.

Loading preview...

Overview

laion/a3-rl-laion_exp_rpt_stack-bash-v3-70-8B is an 8 billion parameter model developed by LAION, originating from an experimental Reinforcement Learning (RL) fine-tuning series. It is based on the laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink (Qwen3-8B) architecture.

Key Characteristics

  • Partial Training: This model represents a partial and cancelled run from the a3 RL series, specifically checkpoint global_step_70.
  • Experimental Artifact: It is preserved as a run artifact, not a fully converged or production-ready model.
  • Low Reward Signal: During its limited training, the model exhibited a low and flat average raw reward (0.08-0.21), indicating non-convergence.
  • Training Data: Fine-tuned on the laion/exp_rpt_stack-bash-v3 agentic task set.

Intended Use

This model is not recommended for general use cases due to its incomplete training and experimental nature. It serves primarily as a historical artifact of a cancelled research run. Researchers interested in the training traces can find companion datasets at DCAgent/a3-rl-laion_exp_rpt_stack-bash-v3.