promotion/qwen3-8b-aaai27-flagship-dpo-s43

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 12, 2026Architecture:Transformer Featherless Exclusive Cold

The promotion/qwen3-8b-aaai27-flagship-dpo-s43 model is an 8 billion parameter language model based on the Qwen3 architecture, fine-tuned using the DPO method. Developed as a research checkpoint for the RONPO AAAI revision experiments, it features a 32K context length. This model is specifically intended for reproducibility and evaluation within the context of the RONPO paper, rather than for general production use.

Loading preview...

Model Overview

The promotion/qwen3-8b-aaai27-flagship-dpo-s43 is an 8 billion parameter language model derived from the Qwen/Qwen3-8B base model. It was fine-tuned using the Direct Preference Optimization (DPO) method with a specific seed (43) and underwent 900 optimizer steps with an effective batch size of 16. This model serves as a research checkpoint for the RONPO AAAI revision experiments.

Key Characteristics

  • Base Model: Qwen3-8B
  • Fine-tuning Method: Direct Preference Optimization (DPO)
  • Context Length: 32,768 tokens
  • Development Context: Part of the AAAI-27 P1 matched budget, passing non-thinking and collapse stability gates.

Intended Use

This model is primarily designed for:

  • Reproducibility: Facilitating the replication of results for the RONPO paper.
  • Evaluation: Serving as a specific checkpoint for research and analysis related to the RONPO paper.

Important Note: This checkpoint is not intended for use as a production assistant or for general-purpose applications. Its utility is confined to the specific research context for which it was developed.