promotion/qwen3-8b-aaai27-flagship-ipo-s43
The promotion/qwen3-8b-aaai27-flagship-ipo-s43 model is an 8 billion parameter research checkpoint based on Qwen/Qwen3-8B, developed by promotion for AAAI revision experiments. This model utilizes the IPO method and is specifically intended for reproducibility and evaluation within the context of the RONPO paper. It is designed for research purposes, focusing on experimental validation rather than production assistant applications.
Loading preview...
Overview
This model, promotion/qwen3-8b-aaai27-flagship-ipo-s43, is an 8 billion parameter research checkpoint derived from the Qwen/Qwen3-8B base model. Developed by promotion, it is a specific iteration from the AAAI revision experiments, utilizing the IPO (Implicit Policy Optimization) method.
Key Characteristics
- Base Model: Qwen/Qwen3-8B
- Methodology: Implicit Policy Optimization (IPO)
- Training Details: Trained with 900 optimizer steps and an effective batch size of 16, matching the AAAI-27 P1 budget.
- Stability: Passed non-thinking and collapse stability gates during its development.
Intended Use
This checkpoint is primarily for reproducibility and evaluation related to the RONPO paper. It serves as an experimental artifact for research validation and is explicitly not intended for use as a production assistant.