laion/a3-rl-DCAgent_mix_h4_binary_easy-50-8B

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 5, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The laion/a3-rl-DCAgent_mix_h4_binary_easy-50-8B is an 8 billion parameter reinforcement learning (RL) model, fine-tuned from a GLM-SFT base. Developed by laion, this model is specifically optimized for the DCAgent/mix_h4_binary_easy task mix, demonstrating a pass@8 score of 0.672 at global_step 50. It is designed for tasks requiring policy learning within a defined environment, leveraging a 32768-token context length.

Loading preview...

Model Overview

The laion/a3-rl-DCAgent_mix_h4_binary_easy-50-8B is an 8 billion parameter reinforcement learning (RL) model, developed by laion. It is a fine-tuned checkpoint from the a3 GLM-SFT base model, specifically laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink.

Key Capabilities & Training Details

  • Task Optimization: This model is trained on the DCAgent/mix_h4_binary_easy task mix, indicating its specialization in specific agent-based decision-making environments.
  • Training Process: The model underwent 2 epochs of training, completing at global_step 67. The selected checkpoint for deployment is at global_step 50, chosen based on the highest trailing-5 EMA of reward/avg_raw_reward (0.4991).
  • Performance Metric: At global_step 50, the model achieved a pass@8 score of 0.672, alongside a raw avg_raw_reward of 0.529.
  • Context Length: It supports a context length of 32768 tokens.

Associated Resources

  • Training Traces: Companion training-time Daytona/Harbor rollouts are available as a dataset: penfever/a3-rl-DCAgent_mix_h4_binary_easy. This dataset contains the last episode of each trial, reflecting the data the policy was trained on.
  • Training Logs: Detailed metrics, reports, and raw chain logs can be found in the training_logs/ directory, providing insights into the training process and performance evolution.