promotion/ronpo-qwen25-1p5b-conflict-matchedbudget-s42

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 12, 2026Architecture:Transformer Featherless Exclusive Cold

This model, promotion/ronpo-qwen25-1p5b-conflict-matchedbudget-s42, is a 1.5 billion parameter language model based on the Qwen2.5-1.5B-Instruct architecture. It was developed as a research checkpoint for RONPO (full objective-response adversary) AAAI revision experiments, specifically a 900-step matched-budget conflict study. The model is intended for reproducibility and evaluation within the context of the RONPO paper, rather than for production use.

Loading preview...

Model Overview

promotion/ronpo-qwen25-1p5b-conflict-matchedbudget-s42 is a 1.5 billion parameter research checkpoint derived from the Qwen/Qwen2.5-1.5B-Instruct base model. It was developed as part of the RONPO (full objective-response adversary) AAAI revision experiments, focusing on a 900-step matched-budget conflict study with a seed of 42. The model's training involved a validation-selected process, optimizing for primary objectives of helpfulness and brevity.

Key Characteristics

  • Base Model: Qwen/Qwen2.5-1.5B-Instruct, a 1.5B parameter instruction-tuned model.
  • Training Method: Utilizes the RONPO full objective-response adversary method.
  • Optimization: Trained with a 900-step matched-budget conflict study, selected based on validation performance.
  • Primary Objectives: Focused on improving helpfulness and brevity in responses.

Performance and Limitations

During sealed 620-prompt testing, the model's helpfulness-brevity prompt-worst score was 0.292996, which did not exceed the strongest baseline (SimPO at 0.315013). The paired difference was -0.022017 (95% CI [-0.055623, 0.011453]), indicating a negative primary result. This checkpoint is released for transparent reproducibility of the RONPO paper's findings, not as a superior performing model.

Intended Use

This model is specifically intended for:

  • Reproducibility: Facilitating the replication of experiments described in the RONPO paper.
  • Evaluation: Serving as a research artifact for further analysis and comparison within academic contexts.

It is explicitly not intended for use as a production assistant or in general-purpose applications.