kaonai/grpo-kaon3-popft-rm600-mouse-reset-timeskip-lr1e4-b004-step125
The kaonai/grpo-kaon3-popft-rm600-mouse-reset-timeskip-lr1e4-b004-step125 is a 26 billion parameter full-weight merged model, based on kaonai/kaon-c-gemma4-26b-v10.1. This model was fine-tuned using a specific GRPO (Generalized Reward Policy Optimization) configuration with a learning rate of 1e-4 and beta of 0.04, utilizing a population-final-transition reward model. It is designed for tasks requiring a robust 26B parameter model with a 32768 token context length, resulting from a precise merging process.
Loading preview...
Model Overview
The kaonai/grpo-kaon3-popft-rm600-mouse-reset-timeskip-lr1e4-b004-step125 is a 26 billion parameter language model, provided as a standalone BF16 full-weight merge. It is not an adapter-only repository, meaning the full model weights are included.
Key Characteristics
- Base Model: Built upon
kaonai/kaon-c-gemma4-26b-v10.1. - Fine-tuning Method: Utilizes a specific Generalized Reward Policy Optimization (GRPO) configuration.
- GRPO Parameters: Features a learning rate of
1e-4and a beta value of0.04. - Reward Model: Incorporates
kaonai/population-final-transition-rm-existing-explicit-s42-step600for its reward signal. - Context Length: Supports a substantial context window of 32768 tokens.
- Merge Details: The model was merged with a
bfloat16dtype, ensuring representative-logit parity was maintained.
Intended Use Cases
This model is suitable for applications requiring a powerful 26B parameter model with a large context window, particularly where the specific GRPO fine-tuning methodology and reward model alignment are beneficial. Its full-weight merge nature makes it ready for direct deployment in various NLP tasks.