promotion/qwen3-8b-aaai27-flagship-ht-mnpo-helpfulness-s42
The promotion/qwen3-8b-aaai27-flagship-ht-mnpo-helpfulness-s42 model is an 8 billion parameter research checkpoint based on Qwen/Qwen3-8B, developed for the RONPO AAAI revision experiments. It utilizes the ht_mnpo_helpfulness method and has a context length of 32768 tokens. This model is specifically intended for reproducibility and evaluation within the RONPO paper, rather than production use.
Loading preview...
Model Overview
This model, promotion/qwen3-8b-aaai27-flagship-ht-mnpo-helpfulness-s42, is an 8 billion parameter research checkpoint derived from the Qwen/Qwen3-8B base model. It was developed as part of the RONPO AAAI revision experiments, specifically utilizing the ht_mnpo_helpfulness method with a seed of 42.
Key Characteristics
- Base Model: Qwen/Qwen3-8B
- Methodology: Implements the
ht_mnpo_helpfulnessmethod. - Training Details: Underwent 900 optimizer steps with an effective batch size of 16.
- Stability: Passed non-thinking and collapse stability gates during its development.
- Context Length: Supports a context length of 32768 tokens.
Intended Use
This checkpoint is primarily intended for reproducibility and evaluation in the context of the RONPO paper. It is explicitly stated that this model is not intended for use as a production assistant.