kaonai/grpo-kaon3-popft-rm600-mouse-reset-timeskip-lr1e4-b004-step175
The kaonai/grpo-kaon3-popft-rm600-mouse-reset-timeskip-lr1e4-b004-step175 is a 26 billion parameter language model, derived from the kaonai/kaon-c-gemma4-26b-v10.1 base model, with a context length of 32768 tokens. This model is a full-weight BF16 merge, specifically fine-tuned using a GRPO (Generalized Reward Policy Optimization) approach with a learning rate of 1e-4 and beta of 0.04. It leverages the kaonai/population-final-transition-rm-existing-explicit-s42-step600 reward model, indicating a specialization in reinforcement learning from human feedback (RLHF) or similar reward-based optimization for specific behavioral alignment.
Loading preview...
Model Overview
The kaonai/grpo-kaon3-popft-rm600-mouse-reset-timeskip-lr1e4-b004-step175 is a 26 billion parameter language model representing a full-weight bfloat16 merge. It is built upon the kaonai/kaon-c-gemma4-26b-v10.1 base architecture, indicating a foundation in the Gemma 4 family of models.
Key Characteristics
- Architecture Base: Derived from
kaonai/kaon-c-gemma4-26b-v10.1. - Parameter Count: 26 billion parameters.
- Merge Type: This is a standalone
BF16full-weight merge, not an adapter-only repository, meaning it contains the complete model weights. - Fine-tuning Method: Utilizes a GRPO (Generalized Reward Policy Optimization) approach with specific hyperparameters: a learning rate of
1e-4and a beta value of0.04. - Reward Model: The fine-tuning process incorporated the
kaonai/population-final-transition-rm-existing-explicit-s42-step600reward model, suggesting a focus on optimizing for specific reward signals or behavioral objectives. - Context Length: Supports a context length of 32768 tokens.
- Integrity Check: The model passed a saved/reloaded representative-logit parity check, ensuring consistency after saving and reloading.
Potential Use Cases
This model is likely suited for applications requiring a model fine-tuned with advanced reinforcement learning techniques, particularly where alignment with a specific reward function is critical. Its full-weight merge nature makes it ready for direct deployment in scenarios benefiting from its specialized GRPO training.