agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp10-q4v3-vs16

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 15, 2026Architecture:Transformer Featherless Exclusive Cold

The agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp10-q4v3-vs16 is a 4 billion parameter Qwen3-4B-Instruct-based model, fine-tuned using OpenRLHF's GRPO reinforcement learning algorithm. This model is specifically optimized for code generation tasks, demonstrating improved performance on problems requiring correct programmatic output. It achieves a pass@8 score of 0.0729 on a held-out validation set of coding challenges, making it suitable for applications demanding high accuracy in generating functional code.

Loading preview...

Model Overview

This model, agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp10-q4v3-vs16, is a 4-billion parameter language model based on Qwen/Qwen3-4B-Instruct-2507. It has been fine-tuned using the OpenRLHF GRPO (Group-Normalized Advantages, No KL Penalty) reinforcement learning algorithm, directly applied to the base Qwen3-4B model without an initial Supervised Fine-Tuning (SFT) seed.

Key Capabilities & Training

  • Code Generation Optimization: The model's primary strength lies in generating correct code, as evidenced by its training objective. It was trained and validated on the "cobalt-train ≤2/64 frontier" dataset, which consists of coding problems that the base model struggled with.
  • Reward Signal: Training utilized a binary code-correctness reward signal, meaning the model received a 1.0 reward if its generated program passed problem tests and 0.0 otherwise.
  • Performance Metrics: At global step 24 of its RL run, the model achieved a pass@8 score of 0.0729 on a held-out validation set, indicating its ability to solve coding problems when allowed multiple attempts.
  • RL Recipe: Key components of the GRPO algorithm included a stop-properly penalty for truncated samples, a DAPO overlong penalty for responses exceeding a token cap, and specific configurations for rollout and training batch sizes.

When to Use This Model

This model is particularly well-suited for use cases requiring robust code generation capabilities, especially for solving programming challenges where functional correctness is paramount. Its optimization for pass@8 suggests it can effectively generate multiple solutions to increase the likelihood of finding a correct one.