agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp10-vs30v11v

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 30, 2026Architecture:Transformer Featherless Exclusive Cold

The agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp10-vs30v11v model is a 4 billion parameter Qwen3-4B-Instruct-based reinforcement learning checkpoint, specifically optimized for code generation tasks. Developed by agurung, this model was trained using the OpenRLHF GRPO algorithm with a binary code-correctness reward signal. It excels at solving programming problems, achieving a pass@8 score of 4.9315 on a specialized code evaluation frontier. This checkpoint is designed for developers requiring a robust code generation model with a 32768 token context length.

Loading preview...

Model Overview

This model, agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp10-vs30v11v, is a 4 billion parameter Qwen3-4B-Instruct-based reinforcement learning (RL) checkpoint. It was developed by agurung using the OpenRLHF GRPO algorithm, with RL applied directly to the base Qwen3-4B model without an initial Supervised Fine-Tuning (SFT) seed. The model is specifically tuned for code generation and problem-solving.

Key Capabilities & Performance

  • Code Generation: Optimized for generating correct code, using a binary code-correctness reward signal during training.
  • Performance: Achieved a pass@8 score of 4.9315 and a pass@1 score of 2.4925 on the cobalt-train <=2/64 frontier evaluation set, indicating strong performance in solving programming challenges.
  • Training Methodology: Utilizes GRPO (group-normalized advantages, no KL penalty) with specific penalties for truncated samples and overlong responses to improve generation quality.

Use Cases

  • Code Problem Solving: Ideal for tasks requiring the generation of functional code that passes specific tests.
  • Programming Assistance: Can be used in applications where robust and accurate code snippets or solutions are needed.

This checkpoint represents the best performing model by pass@8 in its training run, making it suitable for developers focused on high-quality code output.