agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp10-base-q4v3
The agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp10-base-q4v3 is a 4 billion parameter OpenRLHF GRPO reinforcement-learning checkpoint based on Qwen3-4B-Instruct-2507, specifically optimized for code generation. This model was seeded directly from the base Qwen3-4B and fine-tuned using binary code-correctness as a reward signal. It achieves a pass@8 score of 0.1244 on the cobalt-train frontier, making it suitable for tasks requiring robust code problem-solving capabilities.
Loading preview...
Model Overview
This model, agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp10-base-q4v3, is a 4 billion parameter OpenRLHF GRPO reinforcement-learning checkpoint derived from Qwen/Qwen3-4B-Instruct-2507. It was developed by agurung and represents a significant step in enhancing code generation capabilities through direct RL application to a base model.
Key Capabilities & Training
- Reinforcement Learning: Utilizes the GRPO algorithm with group-normalized advantages and no KL penalty, focusing on improving performance through iterative learning.
- Code-Correctness Optimization: Trained with a binary code-correctness reward signal, meaning it receives a positive reward only if the generated program passes problem-specific tests.
- Anti-Truncation Shaping: Incorporates a stop-properly penalty where truncated samples receive a -1.0 reward, and an overlong penalty for responses in the last 1024 tokens, ramping up to -0.25.
- Performance: Achieved a pass@8 score of 0.1244 on the
cobalt-trainfrontier, indicating its ability to solve coding problems when given multiple attempts.
Use Cases
This model is particularly well-suited for:
- Automated Code Generation: Generating functional code snippets or solutions for programming challenges.
- Code Problem Solving: Tasks where the correctness of the generated code can be objectively verified through test cases.
- Research in RL for Code: As a strong baseline or checkpoint for further research into reinforcement learning applications in code generation.