kaonai/grpo-kaon3-popft-rm600-mouse-reset-timeskip-lr1e4-b004-step175

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:2Model Size:26BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 3, 2026Architecture:Transformer Featherless Exclusive Cold

The kaonai/grpo-kaon3-popft-rm600-mouse-reset-timeskip-lr1e4-b004-step175 is a 26 billion parameter language model, derived from the kaonai/kaon-c-gemma4-26b-v10.1 base model, with a context length of 32768 tokens. This model is a full-weight BF16 merge, specifically fine-tuned using a GRPO (Generalized Reward Policy Optimization) approach with a learning rate of 1e-4 and beta of 0.04. It leverages the kaonai/population-final-transition-rm-existing-explicit-s42-step600 reward model, indicating a specialization in reinforcement learning from human feedback (RLHF) or similar reward-based optimization for specific behavioral alignment.

Loading preview...

Model Overview

The kaonai/grpo-kaon3-popft-rm600-mouse-reset-timeskip-lr1e4-b004-step175 is a 26 billion parameter language model representing a full-weight bfloat16 merge. It is built upon the kaonai/kaon-c-gemma4-26b-v10.1 base architecture, indicating a foundation in the Gemma 4 family of models.

Key Characteristics

  • Architecture Base: Derived from kaonai/kaon-c-gemma4-26b-v10.1.
  • Parameter Count: 26 billion parameters.
  • Merge Type: This is a standalone BF16 full-weight merge, not an adapter-only repository, meaning it contains the complete model weights.
  • Fine-tuning Method: Utilizes a GRPO (Generalized Reward Policy Optimization) approach with specific hyperparameters: a learning rate of 1e-4 and a beta value of 0.04.
  • Reward Model: The fine-tuning process incorporated the kaonai/population-final-transition-rm-existing-explicit-s42-step600 reward model, suggesting a focus on optimizing for specific reward signals or behavioral objectives.
  • Context Length: Supports a context length of 32768 tokens.
  • Integrity Check: The model passed a saved/reloaded representative-logit parity check, ensuring consistency after saving and reloading.

Potential Use Cases

This model is likely suited for applications requiring a model fine-tuned with advanced reinforcement learning techniques, particularly where alignment with a specific reward function is critical. Its full-weight merge nature makes it ready for direct deployment in scenarios benefiting from its specialized GRPO training.