laion/a3-rl-laion_nemotron-gym-agent-workplace-v2-5-8B

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 2, 2026Architecture:Transformer Featherless Exclusive Cold

The laion/a3-rl-laion_nemotron-gym-agent-workplace-v2-5-8B is an 8 billion parameter reinforcement learning checkpoint, based on the Qwen3-8B architecture. It was trained by laion on the nemotron-gym-agent-workplace-v2 dataset using GRPO/rloo_n methods. This model represents a small, data-limited ablation run, primarily serving as a weak data point for a3 series research rather than a robust general-purpose agent.

Loading preview...

Model Overview

This model, a3-rl-laion_nemotron-gym-agent-workplace-v2-5-8B, is an 8 billion parameter reinforcement learning (RL) checkpoint developed by laion. It is built upon the Qwen3-8B architecture and was trained using GRPO / rloo_n methods. The training utilized the open-athena/nemotron-gym-agent-workplace-v2 dataset.

Key Characteristics

  • Architecture: Qwen3-8B.
  • Training Method: Reinforcement Learning (GRPO / rloo_n).
  • Dataset: open-athena/nemotron-gym-agent-workplace-v2, which consists of only 297 tasks.
  • Checkpoint Selection: global_step_5 was chosen based on the highest 5-period Exponential Moving Average (EMA) of reward/avg_raw_reward (EMA@5 = 0.3987).
  • Purpose: This model is explicitly noted as a weak a3 ablation data point due to the limited dataset and training budget, not intended as a strong, general-purpose agent.

Training Details

The training process was a small, data-limited run, with the dataset being exhausted after approximately 2 epochs. Companion training traces, including the last episode of each trial, are available as a separate dataset: open-athena/a3-rl-laion_nemotron-gym-agent-workplace-v2.