agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp10-q4v3-groot16
agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp10-q4v3-groot16 is a 4 billion parameter Qwen3-based model, fine-tuned using OpenRLHF GRPO for improved code generation. This model is an RL checkpoint specifically optimized for binary code-correctness, achieving a pass@8 score of 0.1244 on held-out validation problems. It excels at generating correct code solutions, particularly for problems where the base Qwen3-4B model struggled.
Loading preview...
Model Overview
This model, agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp10-q4v3-groot16, is a 4 billion parameter Qwen3-based language model. It has been fine-tuned using OpenRLHF GRPO (Group-Normalized Advantages, no KL penalty) and is an RL checkpoint from global step 16 of its training run. The base model used was Qwen/Qwen3-4B-Instruct-2507, with RL applied directly to the base Qwen3-4B without an initial Supervised Fine-Tuning (SFT) seed.
Key Capabilities & Training
- Code Generation Optimization: The model is specifically trained and validated on the "cobalt-train \u22642/64 frontier," focusing on problems the base model solved on at most 2 of 64 samples.
- Reward Signal: Training uses a binary code-correctness reward signal, meaning 1.0 for a generated program that passes problem tests and 0.0 otherwise.
- Performance: Achieved a pass@8 score of 0.1244 on held-out validation problems, making it the best checkpoint by pass@8 in its run.
- RL Recipe: Utilizes GRPO with a stop-properly penalty (-1.0 for truncated samples) and a DAPO overlong penalty ramping to -0.25 for responses in the last 1024 tokens.
When to Use This Model
This model is particularly well-suited for:
- Code generation tasks where correctness is paramount.
- Improving solutions for programming problems that other models might struggle with.
- Research into Reinforcement Learning from Human Feedback (RLHF) applied to code generation.