agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-q4v2-iid16

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 9, 2026Architecture:Transformer Featherless Exclusive Cold

The agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-q4v2-iid16 model is a 4 billion parameter language model developed by agurung, specifically fine-tuned using Reinforcement Learning (RL) for next-chapter planning tasks. It utilizes a continuous full contrastive reward system, trained on 7,075 examples with a 32,768 token context length. This model excels at generating coherent and contextually relevant next-chapter plans for narrative content.

Loading preview...

Model Overview

The agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-q4v2-iid16 is a 4 billion parameter model specifically fine-tuned for next-chapter planning (NCP) tasks. It leverages Reinforcement Learning (RL) with a unique validation-selected checkpoint, chosen at step 88 from RL seed 42.

Key Capabilities

  • Next-Chapter Planning: Optimized to generate coherent and contextually relevant plans for subsequent chapters in a narrative.
  • RL-Tuned: Trained on 7,075 next-chapter planning examples using a GRPO recipe over two episodes, generating up to 4,096 tokens per plan.
  • Advanced Reward System: Selection was based on 1,120 held-out validation examples, utilizing a mean best-of-eight continuous full contrastive reward rather than a simple binary correctness rate. The reward function incorporates ratio-space own improvement, with penalties for other-book/same-book foils.
  • Robust Selection: The selection process ensures no test examples influenced the model's checkpoint, focusing purely on validation performance.
  • Context Handling: Features ring attention and a multiplicative overlong penalty up to 25% for generated tokens, supporting a 32,768 token context length.

Good For

  • Automated Story Generation: Assisting writers or AI systems in outlining future narrative segments.
  • Content Planning: Generating structured plans for sequential content creation.
  • Research in RL for Narrative: Exploring advanced reward mechanisms and RL applications in creative text generation.