promotion/qwen3-8b-aaai27-flagship-sppo-avg-s42
The promotion/qwen3-8b-aaai27-flagship-sppo-avg-s42 is an 8 billion parameter research checkpoint based on the Qwen3-8B model, developed by Qwen. This model was fine-tuned using the sppo_avg method for reproducibility and evaluation in the RONPO paper, specifically for AAAI-27 revision experiments. It features a 32768 token context length and is intended for research purposes rather than production use.
Loading preview...
Model Overview
This model, promotion/qwen3-8b-aaai27-flagship-sppo-avg-s42, is an 8 billion parameter research checkpoint derived from the Qwen/Qwen3-8B base model. It was developed as part of the RONPO paper's AAAI-27 revision experiments.
Key Characteristics
- Methodology: Fine-tuned using the
sppo_avgmethod with a seed of 42. - Training Details: Underwent 900 optimizer steps with an effective batch size of 16, aligning with the AAAI-27 P1 matched budget.
- Stability: Passed non-thinking and collapse stability gates during its development.
- Context Length: Supports a context window of 32768 tokens.
Intended Use
This checkpoint is specifically designed for:
- Reproducibility: Facilitating the replication of experimental results for the RONPO paper.
- Evaluation: Serving as a specific point for evaluation within the research context.
Note: This model is explicitly not intended for use as a production assistant but rather for academic and research evaluation.