promotion/ronpo-qwen25-1p5b-conflict-matchedbudget-s42
This model, promotion/ronpo-qwen25-1p5b-conflict-matchedbudget-s42, is a 1.5 billion parameter language model based on the Qwen2.5-1.5B-Instruct architecture. It was developed as a research checkpoint for RONPO (full objective-response adversary) AAAI revision experiments, specifically a 900-step matched-budget conflict study. The model is intended for reproducibility and evaluation within the context of the RONPO paper, rather than for production use.
Loading preview...
Model Overview
promotion/ronpo-qwen25-1p5b-conflict-matchedbudget-s42 is a 1.5 billion parameter research checkpoint derived from the Qwen/Qwen2.5-1.5B-Instruct base model. It was developed as part of the RONPO (full objective-response adversary) AAAI revision experiments, focusing on a 900-step matched-budget conflict study with a seed of 42. The model's training involved a validation-selected process, optimizing for primary objectives of helpfulness and brevity.
Key Characteristics
- Base Model: Qwen/Qwen2.5-1.5B-Instruct, a 1.5B parameter instruction-tuned model.
- Training Method: Utilizes the RONPO full objective-response adversary method.
- Optimization: Trained with a 900-step matched-budget conflict study, selected based on validation performance.
- Primary Objectives: Focused on improving helpfulness and brevity in responses.
Performance and Limitations
During sealed 620-prompt testing, the model's helpfulness-brevity prompt-worst score was 0.292996, which did not exceed the strongest baseline (SimPO at 0.315013). The paired difference was -0.022017 (95% CI [-0.055623, 0.011453]), indicating a negative primary result. This checkpoint is released for transparent reproducibility of the RONPO paper's findings, not as a superior performing model.
Intended Use
This model is specifically intended for:
- Reproducibility: Facilitating the replication of experiments described in the RONPO paper.
- Evaluation: Serving as a research artifact for further analysis and comparison within academic contexts.
It is explicitly not intended for use as a production assistant or in general-purpose applications.