kaonai/grpo-kaon3-popft-rm600-mouse-reset-timeskip-lr1e4-b004-step125

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:2Model Size:26BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 3, 2026Architecture:Transformer Featherless Exclusive Cold

The kaonai/grpo-kaon3-popft-rm600-mouse-reset-timeskip-lr1e4-b004-step125 is a 26 billion parameter full-weight merged model, based on kaonai/kaon-c-gemma4-26b-v10.1. This model was fine-tuned using a specific GRPO (Generalized Reward Policy Optimization) configuration with a learning rate of 1e-4 and beta of 0.04, utilizing a population-final-transition reward model. It is designed for tasks requiring a robust 26B parameter model with a 32768 token context length, resulting from a precise merging process.

Loading preview...

Model Overview

The kaonai/grpo-kaon3-popft-rm600-mouse-reset-timeskip-lr1e4-b004-step125 is a 26 billion parameter language model, provided as a standalone BF16 full-weight merge. It is not an adapter-only repository, meaning the full model weights are included.

Key Characteristics

  • Base Model: Built upon kaonai/kaon-c-gemma4-26b-v10.1.
  • Fine-tuning Method: Utilizes a specific Generalized Reward Policy Optimization (GRPO) configuration.
  • GRPO Parameters: Features a learning rate of 1e-4 and a beta value of 0.04.
  • Reward Model: Incorporates kaonai/population-final-transition-rm-existing-explicit-s42-step600 for its reward signal.
  • Context Length: Supports a substantial context window of 32768 tokens.
  • Merge Details: The model was merged with a bfloat16 dtype, ensuring representative-logit parity was maintained.

Intended Use Cases

This model is suitable for applications requiring a powerful 26B parameter model with a large context window, particularly where the specific GRPO fine-tuning methodology and reward model alignment are beneficial. Its full-weight merge nature makes it ready for direct deployment in various NLP tasks.