laion/a3-rl-DCAgent_mix_h2_language_proportional-65-8B
laion/a3-rl-DCAgent_mix_h2_language_proportional-65-8B is an 8 billion parameter language model, fine-tuned using Reinforcement Learning (SkyRL/GRPO) on a Qwen3-8B base. Developed by laion, this model is specifically optimized for the DCAgent/mix_h2_language_proportional task mixture. It achieved an EMA reward of 0.4924 and a pass@8 score of 0.7344 at global step 65, indicating strong performance in its specialized domain. This model is designed for tasks requiring robust RL-driven language capabilities.
Loading preview...
Overview
laion/a3-rl-DCAgent_mix_h2_language_proportional-65-8B is an 8 billion parameter language model that has undergone Reinforcement Learning (RL) fine-tuning. It is based on the a3 GLM-SFT Qwen3-8B model and was fine-tuned using the SkyRL/GRPO method.
Key Capabilities
- RL-Optimized Performance: Specifically fine-tuned on the
DCAgent/mix_h2_language_proportionaltask mixture, indicating specialized performance in this domain. - Performance Metrics: At
global_step_65, the model achieved an EMA(reward) of 0.4924 and a raw reward of 0.6816. It also demonstrated apass@8score of 0.7344. - Checkpoint Selection: This particular checkpoint (
global_step_65) was selected as the EMA-best (α=1/3, 5-period EMA overreward/avg_raw_reward) within the intended run steps.
Training Details
Training traces, including the last episode of each trial, are available as a companion dataset: penfever/a3-rl-DCAgent_mix_h2_language_proportional. These traces represent the rollouts the policy was trained on after rollback/truncation.