kaonai/grpo-kaon3-popft-rm600-mouse-reset-timeskip-lr1e4-b004-step75

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:2Model Size:26BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 3, 2026Architecture:Transformer Featherless Exclusive Cold

The kaonai/grpo-kaon3-popft-rm600-mouse-reset-timeskip-lr1e4-b004-step75 is a 26 billion parameter full-weight merged model based on kaonai/kaon-c-gemma4-26b-v10.1. This model was developed by kaonai using a GRPO (Generalized Reward Policy Optimization) fine-tuning approach with a specific reward model. It is designed for tasks requiring a robust language model with a 32K context length, resulting from a detailed population final-transition optimization process.

Loading preview...

Model Overview

This model, kaonai/grpo-kaon3-popft-rm600-mouse-reset-timeskip-lr1e4-b004-step75, is a 26 billion parameter full-weight merged checkpoint, not an adapter-only repository. It is built upon the kaonai/kaon-c-gemma4-26b-v10.1 base model and has undergone a specific fine-tuning process.

Key Characteristics

  • Base Model: Derived from kaonai/kaon-c-gemma4-26b-v10.1.
  • Fine-tuning Method: Utilizes a GRPO (Generalized Reward Policy Optimization) approach with a learning rate of 1e-4 and a beta value of 0.04.
  • Reward Model: The fine-tuning process incorporated kaonai/population-final-transition-rm-existing-explicit-s42-step600 as its reward model.
  • Merge Type: The model is a bfloat16 full-weight merge, ensuring high precision.
  • Verification: It passed representative-logit parity checks upon saving and reloading, indicating consistency and integrity.

Potential Use Cases

This model is suitable for applications requiring a powerful 26B parameter language model that has been specifically optimized through a reward-based fine-tuning process. Its full-weight merge and bfloat16 precision suggest it can handle complex language generation and understanding tasks where the specific GRPO optimization is beneficial.