agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp10-vs30v11v
The agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp10-vs30v11v model is a 4 billion parameter Qwen3-4B-Instruct-based reinforcement learning checkpoint, specifically optimized for code generation tasks. Developed by agurung, this model was trained using the OpenRLHF GRPO algorithm with a binary code-correctness reward signal. It excels at solving programming problems, achieving a pass@8 score of 4.9315 on a specialized code evaluation frontier. This checkpoint is designed for developers requiring a robust code generation model with a 32768 token context length.
Loading preview...
Model Overview
This model, agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp10-vs30v11v, is a 4 billion parameter Qwen3-4B-Instruct-based reinforcement learning (RL) checkpoint. It was developed by agurung using the OpenRLHF GRPO algorithm, with RL applied directly to the base Qwen3-4B model without an initial Supervised Fine-Tuning (SFT) seed. The model is specifically tuned for code generation and problem-solving.
Key Capabilities & Performance
- Code Generation: Optimized for generating correct code, using a binary code-correctness reward signal during training.
- Performance: Achieved a pass@8 score of 4.9315 and a pass@1 score of 2.4925 on the
cobalt-train <=2/64 frontierevaluation set, indicating strong performance in solving programming challenges. - Training Methodology: Utilizes GRPO (group-normalized advantages, no KL penalty) with specific penalties for truncated samples and overlong responses to improve generation quality.
Use Cases
- Code Problem Solving: Ideal for tasks requiring the generation of functional code that passes specific tests.
- Programming Assistance: Can be used in applications where robust and accurate code snippets or solutions are needed.
This checkpoint represents the best performing model by pass@8 in its training run, making it suitable for developers focused on high-quality code output.