laion/a3-rl-DCAgent_mix_h2_language_proportional-65-8B

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 5, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

laion/a3-rl-DCAgent_mix_h2_language_proportional-65-8B is an 8 billion parameter language model, fine-tuned using Reinforcement Learning (SkyRL/GRPO) on a Qwen3-8B base. Developed by laion, this model is specifically optimized for the DCAgent/mix_h2_language_proportional task mixture. It achieved an EMA reward of 0.4924 and a pass@8 score of 0.7344 at global step 65, indicating strong performance in its specialized domain. This model is designed for tasks requiring robust RL-driven language capabilities.

Loading preview...

Overview

laion/a3-rl-DCAgent_mix_h2_language_proportional-65-8B is an 8 billion parameter language model that has undergone Reinforcement Learning (RL) fine-tuning. It is based on the a3 GLM-SFT Qwen3-8B model and was fine-tuned using the SkyRL/GRPO method.

Key Capabilities

  • RL-Optimized Performance: Specifically fine-tuned on the DCAgent/mix_h2_language_proportional task mixture, indicating specialized performance in this domain.
  • Performance Metrics: At global_step_65, the model achieved an EMA(reward) of 0.4924 and a raw reward of 0.6816. It also demonstrated a pass@8 score of 0.7344.
  • Checkpoint Selection: This particular checkpoint (global_step_65) was selected as the EMA-best (α=1/3, 5-period EMA over reward/avg_raw_reward) within the intended run steps.

Training Details

Training traces, including the last episode of each trial, are available as a companion dataset: penfever/a3-rl-DCAgent_mix_h2_language_proportional. These traces represent the rollouts the policy was trained on after rollback/truncation.