promotion/qwen3-8b-kto-avg-beta0p01-s42
The promotion/qwen3-8b-kto-avg-beta0p01-s42 is a research checkpoint based on the Qwen3-8B model, fine-tuned using the KTO-avg method with a beta of 0.01 and seed 42. This model is specifically intended for reproducibility and evaluation within the context of RONPO AAAI revision experiments. It is not designed for production use but serves as a baseline for research purposes.
Loading preview...
Overview
This model, promotion/qwen3-8b-kto-avg-beta0p01-s42, is a research checkpoint derived from the Qwen/Qwen3-8B base model. It was developed for the RONPO AAAI revision experiments, focusing on specific fine-tuning methodologies.
Key Characteristics
- Methodology: Fine-tuned using the KTO-avg method with a beta value of 0.01.
- Base Model: Built upon the
Qwen3-8Barchitecture, utilizing a non-thinking generation protocol. - Seed: Training was conducted with a seed of 42 for reproducibility.
- Purpose: Primarily intended for research reproducibility and evaluation within the scope of the RONPO paper.
Intended Use
This checkpoint is specifically for academic and research purposes, particularly for evaluating the KTO baseline on an averaged three-reward oracle. It is not recommended for use as a production assistant due to its experimental nature and specific research focus.