agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp10-q4v3-vs16
The agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp10-q4v3-vs16 is a 4 billion parameter Qwen3-4B-Instruct-based model, fine-tuned using OpenRLHF's GRPO reinforcement learning algorithm. This model is specifically optimized for code generation tasks, demonstrating improved performance on problems requiring correct programmatic output. It achieves a pass@8 score of 0.0729 on a held-out validation set of coding challenges, making it suitable for applications demanding high accuracy in generating functional code.
Loading preview...
Model Overview
This model, agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp10-q4v3-vs16, is a 4-billion parameter language model based on Qwen/Qwen3-4B-Instruct-2507. It has been fine-tuned using the OpenRLHF GRPO (Group-Normalized Advantages, No KL Penalty) reinforcement learning algorithm, directly applied to the base Qwen3-4B model without an initial Supervised Fine-Tuning (SFT) seed.
Key Capabilities & Training
- Code Generation Optimization: The model's primary strength lies in generating correct code, as evidenced by its training objective. It was trained and validated on the "cobalt-train ≤2/64 frontier" dataset, which consists of coding problems that the base model struggled with.
- Reward Signal: Training utilized a binary code-correctness reward signal, meaning the model received a 1.0 reward if its generated program passed problem tests and 0.0 otherwise.
- Performance Metrics: At global step 24 of its RL run, the model achieved a pass@8 score of 0.0729 on a held-out validation set, indicating its ability to solve coding problems when allowed multiple attempts.
- RL Recipe: Key components of the GRPO algorithm included a stop-properly penalty for truncated samples, a DAPO overlong penalty for responses exceeding a token cap, and specific configurations for rollout and training batch sizes.
When to Use This Model
This model is particularly well-suited for use cases requiring robust code generation capabilities, especially for solving programming challenges where functional correctness is paramount. Its optimization for pass@8 suggests it can effectively generate multiple solutions to increase the likelihood of finding a correct one.