kaonai/grpo-kaon3-population-final-transition-rm600-mouse-original-reset-matched-lr1e4-b004-step125

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:2Model Size:26BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 2, 2026Architecture:Transformer Featherless Exclusive Cold

The kaonai/grpo-kaon3-population-final-transition-rm600-mouse-original-reset-matched-lr1e4-b004-step125 is a 26 billion parameter language model, based on the kaon-c-gemma4-26b-v10.1 architecture. This model is a standalone BF16 full-weight merge, not an adapter-only repository, and is specifically fine-tuned using a GRPO (Generalized Reinforcement Learning from Human Feedback) approach with a learning rate of 1e-4 and beta of 0.04. It is designed for tasks requiring a robust, fine-tuned model derived from a specific reward model, ensuring representative-logit parity.

Loading preview...

Overview

This model, kaonai/grpo-kaon3-population-final-transition-rm600-mouse-original-reset-matched-lr1e4-b004-step125, is a 26 billion parameter language model. It represents a standalone BF16 full-weight merge of checkpoint-125, distinguishing it from adapter-only repositories. The model is built upon the kaonai/kaon-c-gemma4-26b-v10.1 base architecture.

Key Characteristics

  • Base Model: Derived from kaonai/kaon-c-gemma4-26b-v10.1.
  • Fine-tuning Method: Utilizes a GRPO (Generalized Reinforcement Learning from Human Feedback) approach with specific hyperparameters: a learning rate of 1e-4 and a beta value of 0.04.
  • Reward Model: Fine-tuning was guided by the kaonai/population-final-transition-rm-existing-explicit-s42-step600 reward model.
  • Merge Type: A full-weight merge performed in bfloat16 precision.
  • Verification: The model passed representative-logit parity checks upon saving and reloading.

Intended Use

This model is suitable for applications requiring a robust, fine-tuned language model that has undergone a specific GRPO-based optimization process. Its full-weight merge ensures comprehensive integration of the fine-tuning, making it a complete and deployable solution for tasks aligned with its training methodology and reward model.