promotion/qwen3-8b-kto-avg-beta0p1-s42
The promotion/qwen3-8b-kto-avg-beta0p1-s42 is an 8 billion parameter language model based on the Qwen3-8B architecture, developed by Qwen. This research checkpoint utilizes the KTO-avg method with a beta of 0.1 and a seed of 42, fine-tuned on an averaged three-reward oracle. It is specifically intended for reproducibility and evaluation within the RONPO AAAI revision experiments, rather than production use.
Loading preview...
Model Overview
The promotion/qwen3-8b-kto-avg-beta0p1-s42 is an 8 billion parameter research checkpoint derived from the Qwen/Qwen3-8B base model. It was developed by Qwen as part of the RONPO AAAI revision experiments.
Key Characteristics
- Methodology: Employs the KTO-avg (Kahneman-Tversky Optimization) method with a beta value of 0.1.
- Base Model: Built upon
Qwen/Qwen3-8B, incorporating a non-thinking generation protocol. - Training: Fine-tuned using an averaged three-reward oracle.
- Context Length: Supports a context length of 32768 tokens.
Intended Use
This model is primarily designed for:
- Reproducibility: Facilitating the replication of experimental results for the RONPO paper.
- Evaluation: Serving as a specific checkpoint for research and evaluation purposes within the AAAI revision experiments.
Note: This checkpoint is not intended for use as a production assistant but rather for academic and research validation.