promotion/qwen3-8b-dpo-avg-beta0p05-s42
This is an 8 billion parameter research checkpoint model, `qwen3-8b-dpo-avg-beta0p05-s42`, developed by promotion based on Qwen/Qwen3-8B. It utilizes a DPO-avg method on an averaged three-reward oracle, specifically for reproducibility and evaluation in RONPO AAAI revision experiments. The model is not intended for general-purpose production use but rather for research and experimental validation.
Loading preview...
Model Overview
qwen3-8b-dpo-avg-beta0p05-s42 is an 8 billion parameter research checkpoint model derived from Qwen/Qwen3-8B. It was developed by promotion as part of the RONPO AAAI revision experiments.
Key Characteristics
- Methodology: Employs a DPO-avg (Direct Preference Optimization with averaging) approach.
- Reward Oracle: Based on an averaged three-reward oracle, constructed using per-prompt min-max normalization over
Skywork/Skywork-Reward-V2-Llama-3.1-8B,Nexusflow/Athene-RM-8B, andRLHFlow/ArmoRM-Llama3-8B-v0.1. - Training Parameters: Trained with a specific seed (42) and a DPO beta value of 0.05.
- Data Split: Utilizes the existing MNPO/RONPO UltraFeedback data split.
Intended Use
This model is specifically designed for:
- Reproducibility: Facilitating the replication of experimental results.
- Evaluation: Serving as a tool for assessment within the context of the RONPO research paper.
Important Note: This model is explicitly not intended for use as a general-purpose production assistant. Its primary purpose is research and experimental validation.