kaonai/grpo-kaon3-population-final-transition-rm600-mouse-original-reset-matched-lr1e4-b004-step125
The kaonai/grpo-kaon3-population-final-transition-rm600-mouse-original-reset-matched-lr1e4-b004-step125 is a 26 billion parameter language model, based on the kaon-c-gemma4-26b-v10.1 architecture. This model is a standalone BF16 full-weight merge, not an adapter-only repository, and is specifically fine-tuned using a GRPO (Generalized Reinforcement Learning from Human Feedback) approach with a learning rate of 1e-4 and beta of 0.04. It is designed for tasks requiring a robust, fine-tuned model derived from a specific reward model, ensuring representative-logit parity.
Loading preview...
Overview
This model, kaonai/grpo-kaon3-population-final-transition-rm600-mouse-original-reset-matched-lr1e4-b004-step125, is a 26 billion parameter language model. It represents a standalone BF16 full-weight merge of checkpoint-125, distinguishing it from adapter-only repositories. The model is built upon the kaonai/kaon-c-gemma4-26b-v10.1 base architecture.
Key Characteristics
- Base Model: Derived from
kaonai/kaon-c-gemma4-26b-v10.1. - Fine-tuning Method: Utilizes a GRPO (Generalized Reinforcement Learning from Human Feedback) approach with specific hyperparameters: a learning rate of
1e-4and a beta value of0.04. - Reward Model: Fine-tuning was guided by the
kaonai/population-final-transition-rm-existing-explicit-s42-step600reward model. - Merge Type: A full-weight merge performed in
bfloat16precision. - Verification: The model passed representative-logit parity checks upon saving and reloading.
Intended Use
This model is suitable for applications requiring a robust, fine-tuned language model that has undergone a specific GRPO-based optimization process. Its full-weight merge ensures comprehensive integration of the fine-tuning, making it a complete and deployable solution for tasks aligned with its training methodology and reward model.