promotion/qwen3-8b-aaai27-flagship-ht-mnpo-safety-s42
The promotion/qwen3-8b-aaai27-flagship-ht-mnpo-safety-s42 is an 8 billion parameter research checkpoint based on the Qwen3-8B model, developed for the RONPO AAAI revision experiments. This model utilizes the ht_mnpo_safety method and has a context length of 32768 tokens. It is specifically designed for reproducibility and evaluation within the RONPO paper's research context, rather than for general production use.
Loading preview...
Model Overview
This model, promotion/qwen3-8b-aaai27-flagship-ht-mnpo-safety-s42, is an 8 billion parameter research checkpoint derived from the Qwen/Qwen3-8B base model. It was developed as part of the RONPO AAAI revision experiments, specifically implementing the ht_mnpo_safety method.
Key Characteristics
- Base Model: Qwen3-8B
- Methodology:
ht_mnpo_safety - Training Details: The model underwent 900 optimizer steps with an effective batch size of 16, matching the AAAI-27 P1 budget. It successfully passed non-thinking and collapse stability gates during its development.
- Context Length: Supports a context length of 32768 tokens.
Intended Use
This checkpoint is primarily intended for reproducibility and evaluation within the scope of the RONPO paper. It serves as a specific research artifact for academic purposes and is not designed or recommended for use as a production assistant or in general-purpose applications. Its utility lies in validating experimental results and methodologies presented in the research.