promotion/qwen3-8b-ipo-avg-beta0p01-s42

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 11, 2026Architecture:Transformer Featherless Exclusive Cold

The promotion/qwen3-8b-ipo-avg-beta0p01-s42 is an 8 billion parameter research checkpoint based on the Qwen3-8B model, developed for reproducibility and evaluation in the context of RONPO AAAI revision experiments. It utilizes the IPO-avg method with a beta of 0.01 and a non-thinking generation protocol. This model is specifically intended for research purposes related to the RONPO paper and is not designed for production assistant use cases.

Loading preview...

Overview

This model, promotion/qwen3-8b-ipo-avg-beta0p01-s42, is an 8 billion parameter research checkpoint derived from the Qwen/Qwen3-8B base model. It was created as part of the RONPO AAAI revision experiments, focusing on reproducibility and evaluation.

Key Characteristics

  • Methodology: Employs the IPO-avg method with a beta value of 0.01.
  • Base Model: Built upon Qwen/Qwen3-8B, utilizing a non-thinking generation protocol.
  • Seed: Training was conducted with a seed of 42.
  • Purpose: Primarily serves as a baseline for the averaged three-reward oracle within the RONPO paper's experimental framework.

Intended Use

This checkpoint is specifically for reproducibility and evaluation related to the RONPO paper. It is not intended for use as a production assistant or in general-purpose applications. Developers should consider this model for academic research, comparative analysis, or replicating the specific experimental conditions described in the RONPO paper.