kaonai/grpo-kaon3-popft-rm600-mouse-reset-timeskip-lr1e4-b004-step75
The kaonai/grpo-kaon3-popft-rm600-mouse-reset-timeskip-lr1e4-b004-step75 is a 26 billion parameter full-weight merged model based on kaonai/kaon-c-gemma4-26b-v10.1. This model was developed by kaonai using a GRPO (Generalized Reward Policy Optimization) fine-tuning approach with a specific reward model. It is designed for tasks requiring a robust language model with a 32K context length, resulting from a detailed population final-transition optimization process.
Loading preview...
Model Overview
This model, kaonai/grpo-kaon3-popft-rm600-mouse-reset-timeskip-lr1e4-b004-step75, is a 26 billion parameter full-weight merged checkpoint, not an adapter-only repository. It is built upon the kaonai/kaon-c-gemma4-26b-v10.1 base model and has undergone a specific fine-tuning process.
Key Characteristics
- Base Model: Derived from
kaonai/kaon-c-gemma4-26b-v10.1. - Fine-tuning Method: Utilizes a GRPO (Generalized Reward Policy Optimization) approach with a learning rate of
1e-4and a beta value of0.04. - Reward Model: The fine-tuning process incorporated
kaonai/population-final-transition-rm-existing-explicit-s42-step600as its reward model. - Merge Type: The model is a
bfloat16full-weight merge, ensuring high precision. - Verification: It passed representative-logit parity checks upon saving and reloading, indicating consistency and integrity.
Potential Use Cases
This model is suitable for applications requiring a powerful 26B parameter language model that has been specifically optimized through a reward-based fine-tuning process. Its full-weight merge and bfloat16 precision suggest it can handle complex language generation and understanding tasks where the specific GRPO optimization is beneficial.