laion/a3-rl-laion_nemotron-gym-instruction-following-structured-75-8B
The laion/a3-rl-laion_nemotron-gym-instruction-following-structured-75-8B is an 8 billion parameter language model, fine-tuned using Reinforcement Learning (RL) with SkyRL on the open-athena/nemotron-gym-instruction-following-structured dataset. Based on the Qwen3-8B architecture, this model is specifically optimized for structured instruction following tasks. It achieves a high average raw reward, making it suitable for applications requiring precise adherence to structured instructions.
Loading preview...
Model Overview
This model, a3-rl-laion_nemotron-gym-instruction-following-structured-75-8B, is an 8 billion parameter language model that has undergone Reinforcement Learning (RL) fine-tuning. It is built upon the laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink base model, which is a variant of Qwen3-8B.
Key Capabilities
- Structured Instruction Following: The model is specifically fine-tuned on the
open-athena/nemotron-gym-instruction-following-structureddataset, making it highly proficient in understanding and executing structured instructions. - RL Optimization: Fine-tuned using SkyRL, the model's checkpoint was selected based on a 5-period EMA of
reward/avg_raw_reward, achieving an EMA of 0.9512 and a step reward of 0.9902 at global_step 75.
Training Details
The training process involved 80 steps, with the final reward around 0.92 and pass@8 around 0.95. Training traces, including the last episode of each trial, are available as a companion dataset: open-athena/a3-rl-laion_nemotron-gym-instruction-following-structured.
Good For
This model is particularly well-suited for use cases that demand precise and reliable execution of structured instructions, where adherence to format and specific commands is critical.