promotion/qwen3-8b-aaai27-flagship-simpo-s43
The promotion/qwen3-8b-aaai27-flagship-simpo-s43 is an 8 billion parameter language model based on the Qwen3-8B architecture, fine-tuned using the SIMPO method. This research checkpoint is specifically designed for reproducibility and evaluation within the RONPO AAAI revision experiments. It is not intended for general production use but serves as a specific experimental artifact for research purposes.
Loading preview...
Model Overview
The promotion/qwen3-8b-aaai27-flagship-simpo-s43 is a research checkpoint derived from the Qwen/Qwen3-8B base model, featuring 8 billion parameters and a 32K context length. It was fine-tuned using the SIMPO method as part of the RONPO AAAI revision experiments.
Key Characteristics
- Base Model: Qwen3-8B
- Fine-tuning Method: SIMPO (Simple Preference Optimization)
- Experimental Context: Developed for the AAAI-27 P1 matched budget, undergoing 900 optimizer steps with an effective batch size of 16.
- Stability: Passed non-thinking and collapse stability gates during its development.
Intended Use
This model checkpoint is primarily intended for reproducibility and evaluation within the scope of the RONPO research paper. It is explicitly stated that this checkpoint is not designed for use as a production assistant but rather as a specific artifact for academic and experimental validation.