promotion/qwen3-8b-inpo-avg-eta0p005-s42

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 10, 2026Architecture:Transformer Featherless Exclusive Cold

The promotion/qwen3-8b-inpo-avg-eta0p005-s42 model is a research checkpoint based on the Qwen3-8B architecture, developed for reproducibility and evaluation within the RONPO AAAI revision experiments. It utilizes the INPO-avg method with an eta of 0.005 and a non-thinking generation protocol. This model is specifically intended for research purposes related to the RONPO paper and is not designed for production assistant use cases.

Loading preview...

Model Overview

This model, promotion/qwen3-8b-inpo-avg-eta0p005-s42, is a research checkpoint derived from the Qwen3-8B base model. It was developed as part of the RONPO AAAI revision experiments, focusing on specific methodological evaluations.

Key Characteristics

  • Methodology: Implements the INPO-avg method with a learning rate (eta) of 0.005.
  • Base Model: Built upon Qwen/Qwen3-8B, utilizing a non-thinking generation protocol.
  • Purpose: Primarily intended for reproducibility and evaluation within the context of the RONPO paper.
  • Development: Represents a specific iteration (seed 42) from the research experiments.

Intended Use

  • Research: Ideal for researchers and developers looking to replicate or evaluate the RONPO paper's findings.
  • Evaluation: Suitable for assessing the performance of the INPO-avg method under specified conditions.

Note: This checkpoint is explicitly not intended as a production assistant and should be used strictly for research and evaluation purposes.