promotion/qwen3-8b-aaai27-flagship-kto-s44
The promotion/qwen3-8b-aaai27-flagship-kto-s44 is an 8 billion parameter research checkpoint based on the Qwen3-8B model, developed for reproducibility and evaluation in the context of RONPO AAAI revision experiments. This model was fine-tuned using the KTO method with a specific seed (44) and is intended for academic research purposes. It is not designed for production use but serves as a specific experimental artifact.
Loading preview...
Overview
This model, promotion/qwen3-8b-aaai27-flagship-kto-s44, is an 8 billion parameter research checkpoint derived from the Qwen/Qwen3-8B base model. It was developed as part of the RONPO AAAI revision experiments, focusing on reproducibility and evaluation.
Key Characteristics
- Base Model: Qwen3-8B
- Fine-tuning Method: KTO (Kahneman-Tversky Optimization)
- Seed: 44, indicating a specific experimental configuration.
- Training Details: The model underwent 900 optimizer steps with an effective batch size of 16, aligning with the AAAI-27 P1 matched budget. It successfully passed non-thinking and collapse stability gates.
- Purpose: Primarily intended for academic reproducibility and evaluation related to the RONPO paper.
Intended Use
This checkpoint is specifically designed for research and evaluation within the context of the RONPO paper. It is not recommended or intended for use as a production assistant or in general-purpose applications. Its value lies in providing a specific, documented experimental artifact for academic study.