promotion/qwen3-8b-inpo-avg-eta0p01-s42
The promotion/qwen3-8b-inpo-avg-eta0p01-s42 model is a research checkpoint derived from the Qwen3-8B base model, developed by Qwen. This model utilizes the INPO-avg method with an eta of 0.01 and a non-thinking generation protocol. It is specifically intended for reproducibility and evaluation within the context of the RONPO paper, rather than for production use.
Loading preview...
Overview
This model, qwen3-8b-inpo-avg-eta0p01-s42, is a research checkpoint developed by Qwen, specifically for the RONPO AAAI revision experiments. It is based on the Qwen/Qwen3-8B model and incorporates the INPO-avg method with an eta of 0.01, using a non-thinking generation protocol. The model was trained with a seed of 42 and is primarily intended for academic reproducibility and evaluation.
Key Characteristics
- Base Model: Derived from
Qwen/Qwen3-8B. - Methodology: Employs the INPO-avg method with a specific learning rate (eta=0.01).
- Generation Protocol: Utilizes a non-thinking generation protocol.
- Purpose: Created as a research artifact for the RONPO paper, focusing on reproducibility and evaluation.
Intended Use
This checkpoint is explicitly designed for:
- Reproducibility: Facilitating the replication of experiments described in the RONPO paper.
- Evaluation: Serving as a baseline for assessing the performance of the INPO method on the averaged three-reward oracle.
Note: This model is not recommended for use as a production assistant or in general-purpose applications.