agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-q4v3-vs16
agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-q4v3-vs16 is a 4 billion parameter Qwen3-4B model fine-tuned using OpenRLHF GRPO for improved code generation. This reinforcement learning checkpoint, developed by agurung, is specifically optimized for solving programming problems, achieving the best pass@8 score in its training run. It was trained on a curated dataset of challenging coding tasks with a binary code-correctness reward signal, making it highly effective for code-related applications.
Loading preview...
Model Overview
This model, agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-q4v3-vs16, is a 4 billion parameter Qwen3-4B variant that has undergone Reinforcement Learning (RL) using the OpenRLHF GRPO algorithm. It is a specific checkpoint (global step 26) from an RL run, identified as the best performer by its pass@8 metric.
Key Capabilities & Training Details
- Base Model:
Qwen/Qwen3-4B-Instruct-2507was used as the foundation, with RL applied directly to the base model without an initial Supervised Fine-Tuning (SFT) seed. - Optimization Focus: The model is specifically trained for code generation and problem-solving, with a reward signal based on binary code-correctness (1.0 for passing problem tests, 0.0 otherwise).
- Training Data: It was trained and validated on the "cobalt-train \u22642/64 frontier" dataset, comprising 1833 training problems and 112 held-out validation problems that the base model struggled with.
- RL Algorithm: Utilizes GRPO (group-normalized advantages, no KL penalty) with specific penalties for truncated samples (-1.0) and overlong responses (ramping additive penalty up to -0.25 for the last 1024 tokens).
- Efficiency: Trained with 8 samples per prompt, a rollout batch size of 128, and a train batch size of 128, over 2 episodes.
When to Use This Model
This model is particularly well-suited for tasks requiring:
- High-quality code generation for programming challenges.
- Applications where code correctness is a primary evaluation metric.
- Scenarios benefiting from a model fine-tuned with reinforcement learning on challenging coding problems.