promotion/qwen3-8b-aaai27-flagship-dpo-s42
The promotion/qwen3-8b-aaai27-flagship-dpo-s42 model is an 8 billion parameter language model developed by promotion, based on the Qwen3-8B architecture. This research checkpoint was fine-tuned using DPO (Direct Preference Optimization) with a specific seed (42) for reproducibility and evaluation in AAAI-27 revision experiments. It is primarily intended for research purposes related to the RONPO paper, focusing on method evaluation rather than production use.
Loading preview...
Overview
The promotion/qwen3-8b-aaai27-flagship-dpo-s42 is an 8 billion parameter language model derived from the Qwen/Qwen3-8B base model. This specific checkpoint represents a research artifact from the AAAI-27 revision experiments, focusing on the RONPO paper.
Key Characteristics
- Base Model: Qwen3-8B architecture.
- Fine-tuning Method: Utilizes Direct Preference Optimization (DPO).
- Training Details: Trained with 900 optimizer steps and an effective batch size of 16, passing non-thinking and collapse stability gates.
- Purpose: Primarily intended for reproducibility and evaluation within the context of the RONPO research paper.
- Context Length: Supports a context length of 32768 tokens.
Intended Use
This model checkpoint is specifically designed for:
- Research Reproducibility: Facilitating the replication of experimental results for the RONPO paper.
- Evaluation: Serving as a specific point for evaluating the DPO method within the AAAI-27 revision framework.
Note: This model is a research checkpoint and is not intended for use as a production assistant.