promotion/qwen3-8b-dpo-avg-beta0p01-s42

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 10, 2026License:otherArchitecture:Transformer Featherless Exclusive Cold

The promotion/qwen3-8b-dpo-avg-beta0p01-s42 model is an 8 billion parameter language model based on the Qwen3-8B architecture, fine-tuned using DPO-avg with a beta of 0.01. This model is a research checkpoint specifically developed for RONPO AAAI revision experiments, utilizing an averaged three-reward oracle for training. Its primary purpose is for reproducibility and evaluation within the context of the RONPO research paper, rather than general-purpose production use.

Loading preview...

Model Overview

This model, promotion/qwen3-8b-dpo-avg-beta0p01-s42, is an 8 billion parameter research checkpoint derived from the Qwen/Qwen3-8B base model. It was fine-tuned using the DPO-avg (Direct Preference Optimization with averaging) method, specifically for RONPO AAAI revision experiments.

Key Characteristics

  • Base Model: Qwen/Qwen3-8B (non-thinking mode).
  • Fine-tuning Method: DPO-avg, with a DPO beta value of 0.01.
  • Training Data: Utilizes the existing MNPO/RONPO UltraFeedback split.
  • Oracle Construction: Employs a per-prompt min-max normalization over three distinct reward models: Skywork/Skywork-Reward-V2-Llama-3.1-8B, Nexusflow/Athene-RM-8B, and RLHFlow/ArmoRM-Llama3-8B-v0.1, followed by an unweighted average.
  • Context Length: Supports a context length of 32768 tokens.

Intended Use

This model is specifically intended for reproducibility and evaluation within the scope of the RONPO research paper. It is not designed or recommended for use as a general-purpose production assistant due to its specialized research nature.