agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp10-base-q4v3

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 15, 2026Architecture:Transformer Featherless Exclusive Cold

The agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp10-base-q4v3 is a 4 billion parameter OpenRLHF GRPO reinforcement-learning checkpoint based on Qwen3-4B-Instruct-2507, specifically optimized for code generation. This model was seeded directly from the base Qwen3-4B and fine-tuned using binary code-correctness as a reward signal. It achieves a pass@8 score of 0.1244 on the cobalt-train frontier, making it suitable for tasks requiring robust code problem-solving capabilities.

Loading preview...

Model Overview

This model, agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp10-base-q4v3, is a 4 billion parameter OpenRLHF GRPO reinforcement-learning checkpoint derived from Qwen/Qwen3-4B-Instruct-2507. It was developed by agurung and represents a significant step in enhancing code generation capabilities through direct RL application to a base model.

Key Capabilities & Training

  • Reinforcement Learning: Utilizes the GRPO algorithm with group-normalized advantages and no KL penalty, focusing on improving performance through iterative learning.
  • Code-Correctness Optimization: Trained with a binary code-correctness reward signal, meaning it receives a positive reward only if the generated program passes problem-specific tests.
  • Anti-Truncation Shaping: Incorporates a stop-properly penalty where truncated samples receive a -1.0 reward, and an overlong penalty for responses in the last 1024 tokens, ramping up to -0.25.
  • Performance: Achieved a pass@8 score of 0.1244 on the cobalt-train frontier, indicating its ability to solve coding problems when given multiple attempts.

Use Cases

This model is particularly well-suited for:

  • Automated Code Generation: Generating functional code snippets or solutions for programming challenges.
  • Code Problem Solving: Tasks where the correctness of the generated code can be objectively verified through test cases.
  • Research in RL for Code: As a strong baseline or checkpoint for further research into reinforcement learning applications in code generation.