promotion/qwen3-8b-dpo-avg-beta0p05-s42

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 10, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

This is an 8 billion parameter research checkpoint model, `qwen3-8b-dpo-avg-beta0p05-s42`, developed by promotion based on Qwen/Qwen3-8B. It utilizes a DPO-avg method on an averaged three-reward oracle, specifically for reproducibility and evaluation in RONPO AAAI revision experiments. The model is not intended for general-purpose production use but rather for research and experimental validation.

Loading preview...

Model Overview

qwen3-8b-dpo-avg-beta0p05-s42 is an 8 billion parameter research checkpoint model derived from Qwen/Qwen3-8B. It was developed by promotion as part of the RONPO AAAI revision experiments.

Key Characteristics

  • Methodology: Employs a DPO-avg (Direct Preference Optimization with averaging) approach.
  • Reward Oracle: Based on an averaged three-reward oracle, constructed using per-prompt min-max normalization over Skywork/Skywork-Reward-V2-Llama-3.1-8B, Nexusflow/Athene-RM-8B, and RLHFlow/ArmoRM-Llama3-8B-v0.1.
  • Training Parameters: Trained with a specific seed (42) and a DPO beta value of 0.05.
  • Data Split: Utilizes the existing MNPO/RONPO UltraFeedback data split.

Intended Use

This model is specifically designed for:

  • Reproducibility: Facilitating the replication of experimental results.
  • Evaluation: Serving as a tool for assessment within the context of the RONPO research paper.

Important Note: This model is explicitly not intended for use as a general-purpose production assistant. Its primary purpose is research and experimental validation.