promotion/qwen3-8b-dpo-avg-beta0p01-s42
The promotion/qwen3-8b-dpo-avg-beta0p01-s42 model is an 8 billion parameter language model based on the Qwen3-8B architecture, fine-tuned using DPO-avg with a beta of 0.01. This model is a research checkpoint specifically developed for RONPO AAAI revision experiments, utilizing an averaged three-reward oracle for training. Its primary purpose is for reproducibility and evaluation within the context of the RONPO research paper, rather than general-purpose production use.
Loading preview...
Model Overview
This model, promotion/qwen3-8b-dpo-avg-beta0p01-s42, is an 8 billion parameter research checkpoint derived from the Qwen/Qwen3-8B base model. It was fine-tuned using the DPO-avg (Direct Preference Optimization with averaging) method, specifically for RONPO AAAI revision experiments.
Key Characteristics
- Base Model: Qwen/Qwen3-8B (non-thinking mode).
- Fine-tuning Method: DPO-avg, with a DPO beta value of 0.01.
- Training Data: Utilizes the existing MNPO/RONPO UltraFeedback split.
- Oracle Construction: Employs a per-prompt min-max normalization over three distinct reward models:
Skywork/Skywork-Reward-V2-Llama-3.1-8B,Nexusflow/Athene-RM-8B, andRLHFlow/ArmoRM-Llama3-8B-v0.1, followed by an unweighted average. - Context Length: Supports a context length of 32768 tokens.
Intended Use
This model is specifically intended for reproducibility and evaluation within the scope of the RONPO research paper. It is not designed or recommended for use as a general-purpose production assistant due to its specialized research nature.