promotion/qwen3-8b-ipo-avg-beta0p1-s42

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 11, 2026Architecture:Transformer Featherless Exclusive Cold

The promotion/qwen3-8b-ipo-avg-beta0p1-s42 model is an 8 billion parameter research checkpoint based on Qwen/Qwen3-8B, utilizing the IPO-avg method with a beta of 0.1 and a seed of 42. It is specifically designed for reproducibility and evaluation within the context of RONPO AAAI revision experiments. This model is optimized for research purposes related to averaged three-reward oracle baselines and is not intended for production use.

Loading preview...

Model Overview

This model, promotion/qwen3-8b-ipo-avg-beta0p1-s42, is an 8 billion parameter research checkpoint derived from the Qwen/Qwen3-8B base model. It was developed as part of the RONPO AAAI revision experiments.

Key Characteristics

  • Methodology: Implements the IPO-avg method with a beta value of 0.1.
  • Base Model: Built upon Qwen/Qwen3-8B, incorporating a non-thinking generation protocol.
  • Configuration: Utilizes a fixed seed of 42 for reproducibility.
  • Purpose: Serves as a baseline for the averaged three-reward oracle within the RONPO research framework.

Intended Use

This checkpoint is specifically for:

  • Reproducibility: Facilitating the replication of experimental results for the RONPO paper.
  • Evaluation: Providing a consistent model for evaluating research hypotheses.

Important Note: This model is a research artifact and is not intended for use as a production assistant or in real-world applications.