agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-q4v3-vs16

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 18, 2026Architecture:Transformer Featherless Exclusive Cold

agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-q4v3-vs16 is a 4 billion parameter Qwen3-4B model fine-tuned using OpenRLHF GRPO for improved code generation. This reinforcement learning checkpoint, developed by agurung, is specifically optimized for solving programming problems, achieving the best pass@8 score in its training run. It was trained on a curated dataset of challenging coding tasks with a binary code-correctness reward signal, making it highly effective for code-related applications.

Loading preview...

Model Overview

This model, agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-q4v3-vs16, is a 4 billion parameter Qwen3-4B variant that has undergone Reinforcement Learning (RL) using the OpenRLHF GRPO algorithm. It is a specific checkpoint (global step 26) from an RL run, identified as the best performer by its pass@8 metric.

Key Capabilities & Training Details

  • Base Model: Qwen/Qwen3-4B-Instruct-2507 was used as the foundation, with RL applied directly to the base model without an initial Supervised Fine-Tuning (SFT) seed.
  • Optimization Focus: The model is specifically trained for code generation and problem-solving, with a reward signal based on binary code-correctness (1.0 for passing problem tests, 0.0 otherwise).
  • Training Data: It was trained and validated on the "cobalt-train \u22642/64 frontier" dataset, comprising 1833 training problems and 112 held-out validation problems that the base model struggled with.
  • RL Algorithm: Utilizes GRPO (group-normalized advantages, no KL penalty) with specific penalties for truncated samples (-1.0) and overlong responses (ramping additive penalty up to -0.25 for the last 1024 tokens).
  • Efficiency: Trained with 8 samples per prompt, a rollout batch size of 128, and a train batch size of 128, over 2 episodes.

When to Use This Model

This model is particularly well-suited for tasks requiring:

  • High-quality code generation for programming challenges.
  • Applications where code correctness is a primary evaluation metric.
  • Scenarios benefiting from a model fine-tuned with reinforcement learning on challenging coding problems.