promotion/qwen3-8b-kto-avg-beta0p05-s42
The promotion/qwen3-8b-kto-avg-beta0p05-s42 model is a research checkpoint based on the Qwen3-8B architecture, developed for reproducibility and evaluation in the context of RONPO AAAI revision experiments. It utilizes the KTO-avg method with a beta of 0.05 and a non-thinking generation protocol. This model is specifically intended for research purposes related to the RONPO paper and is not designed for production assistant use cases.
Loading preview...
Model Overview
This model, qwen3-8b-kto-avg-beta0p05-s42, is a research checkpoint derived from the Qwen/Qwen3-8B base model. It was developed as part of the RONPO AAAI revision experiments, focusing on the KTO-avg method with a beta value of 0.05 and a non-thinking generation protocol. The model was trained with a specific seed (42) and represents a KTO baseline on an averaged three-reward oracle.
Key Characteristics
- Base Model: Qwen3-8B
- Training Method: KTO-avg with beta=0.05
- Generation Protocol: Non-thinking
- Purpose: Research checkpoint for reproducibility and evaluation of the RONPO paper.
Intended Use
This checkpoint is primarily intended for:
- Reproducibility: Facilitating the replication of experimental results for the RONPO paper.
- Evaluation: Serving as a baseline for assessing the performance of the KTO method within the RONPO research context.
Important Note: This model is explicitly stated as not intended as a production assistant and should be used solely for its designated research and evaluation purposes.