promotion/qwen3-8b-inpo-avg-eta0p0075-s42

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 10, 2026Architecture:Transformer Featherless Exclusive Cold

This model is a research checkpoint, `qwen3-8b-inpo-avg-eta0p0075-s42`, based on the Qwen/Qwen3-8B architecture. It was developed using the INPO-avg method with an eta of 0.0075 and a seed of 42, specifically for the RONPO AAAI revision experiments. Its primary purpose is for reproducibility and evaluation within the context of the RONPO paper, rather than for production assistant applications.

Loading preview...

Overview

This model, qwen3-8b-inpo-avg-eta0p0075-s42, is a research checkpoint derived from the Qwen/Qwen3-8B base model. It was developed as part of the RONPO AAAI revision experiments, utilizing the INPO-avg method with a specific eta value of 0.0075 and a seed of 42. The training involved a non-thinking generation protocol.

Key Characteristics

  • Base Model: Qwen/Qwen3-8B
  • Methodology: INPO-avg with eta=0.0075
  • Seed: 42
  • Purpose: Primarily for reproducibility and evaluation related to the RONPO paper.

Intended Use

This checkpoint serves as an INPO baseline on an averaged three-reward oracle. It is explicitly stated that this model is not intended as a production assistant but rather for research and evaluation purposes to support the RONPO paper.