laion/a3-rl-laion_nemotron-gym-agent-workplace-v2-5-8B
The laion/a3-rl-laion_nemotron-gym-agent-workplace-v2-5-8B is an 8 billion parameter reinforcement learning checkpoint, based on the Qwen3-8B architecture. It was trained by laion on the nemotron-gym-agent-workplace-v2 dataset using GRPO/rloo_n methods. This model represents a small, data-limited ablation run, primarily serving as a weak data point for a3 series research rather than a robust general-purpose agent.
Loading preview...
Model Overview
This model, a3-rl-laion_nemotron-gym-agent-workplace-v2-5-8B, is an 8 billion parameter reinforcement learning (RL) checkpoint developed by laion. It is built upon the Qwen3-8B architecture and was trained using GRPO / rloo_n methods. The training utilized the open-athena/nemotron-gym-agent-workplace-v2 dataset.
Key Characteristics
- Architecture: Qwen3-8B.
- Training Method: Reinforcement Learning (GRPO / rloo_n).
- Dataset:
open-athena/nemotron-gym-agent-workplace-v2, which consists of only 297 tasks. - Checkpoint Selection:
global_step_5was chosen based on the highest 5-period Exponential Moving Average (EMA) ofreward/avg_raw_reward(EMA@5 = 0.3987). - Purpose: This model is explicitly noted as a weak a3 ablation data point due to the limited dataset and training budget, not intended as a strong, general-purpose agent.
Training Details
The training process was a small, data-limited run, with the dataset being exhausted after approximately 2 epochs. Companion training traces, including the last episode of each trial, are available as a separate dataset: open-athena/a3-rl-laion_nemotron-gym-agent-workplace-v2.