agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-q4v2-vs16

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 9, 2026Architecture:Transformer Featherless Exclusive Cold

The agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-q4v2-vs16 is a 4 billion parameter model specifically trained for next-chapter planning. It utilizes a GRPO-based Reinforcement Learning approach, fine-tuned on 7,075 planning examples and validated using a continuous full contrastive reward system. This model is distinguished by its unique reward function, which incorporates ratio-space own improvement and foil penalties, making it suitable for generating coherent and contextually relevant chapter continuations.

Loading preview...

Model Overview

The agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-q4v2-vs16 is a 4 billion parameter model developed by agurung, specifically engineered for next-chapter planning (NCP) tasks. This model is a validation-selected Reinforcement Learning (RL) checkpoint, identified as Arm vs16, from RL seed 42 at step 64.

Key Training and Selection Details

  • Training Data: The model was trained on 7,075 next-chapter planning examples.
  • Validation: Selection was performed using 1,120 held-out validation examples, with eight samples per example. The selection metric is a mean best-of-eight continuous full contrastive reward, rather than a simple binary correctness rate, ensuring a nuanced evaluation of plan quality.
  • Reward Function: A distinctive reward mechanism is employed:
    • Ratio-space own improvement above 5.
    • Penalties for other-book/same-book foils above 15, with weights of 0.5/0.25 respectively.
    • Own reward is uncapped, while malformed plans incur a -25 penalty.
  • RL Recipe: The training utilized a GRPO (Generalized Reinforcement Policy Optimization) approach over two episodes, generating 4,096 tokens per sample. It includes a multiplicative overlong penalty up to 25%, a truncated-sample override of -1, and ring attention 2.

What Makes This Model Different?

This model stands out due to its specialized training for next-chapter planning and its sophisticated reward function. Unlike models optimized for general text generation, this model's reward system directly targets the coherence and contextual relevance of chapter plans, penalizing inconsistencies and rewarding meaningful progression. The use of continuous full contrastive reward for validation provides a more robust selection criterion than typical binary metrics.

Use Cases

This model is particularly well-suited for applications requiring:

  • Automated story generation: Generating logical and engaging continuations for narratives.
  • Creative writing assistance: Aiding authors in outlining plot points and chapter structures.
  • Interactive fiction: Powering AI agents that can guide story progression based on user input.