agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-base-q4v2

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 9, 2026Architecture:Transformer Featherless Exclusive Cold

The agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-base-q4v2 model is a 4 billion parameter language model developed by agurung, fine-tuned using Reinforcement Learning (RL) with GRPO. It is specifically optimized for next-chapter planning tasks, trained on 7,075 examples and selected based on a continuous full contrastive reward system. This model excels at generating coherent and contextually relevant next-chapter plans, distinguishing it from general-purpose LLMs.

Loading preview...

Model Overview

The agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-base-q4v2 is a 4 billion parameter language model that has undergone Reinforcement Learning (RL) fine-tuning. Developed by agurung, this model is specifically designed for next-chapter planning tasks, leveraging a GRPO (Generalized Reinforcement Learning with Policy Optimization) recipe.

Key Capabilities

  • Next-Chapter Planning: The model is trained on 7,075 next-chapter planning examples, focusing on generating coherent and contextually appropriate continuations.
  • RL-based Optimization: It utilizes a unique continuous full contrastive reward system for selection, based on 1,120 held-out validation examples, rather than a simple binary correctness rate. This reward system measures the model's own improvement against foil penalties.
  • Advanced Training Methodology: The training involved two episodes with 4,096 generated tokens, incorporating a multiplicative overlong penalty and ring attention 2.

Good For

  • Automated Story Generation: Ideal for applications requiring the generation of logical and engaging next chapters or plot points in narrative content.
  • Content Planning Tools: Suitable for tools that assist writers or creators in structuring long-form content by suggesting subsequent sections or developments.
  • Research in RL for Text Generation: Provides a specialized model for exploring advanced RL techniques in text generation, particularly for structured output like planning.