promotion/ronpo-llama31-saferlhf-stage3-os-s42
The promotion/ronpo-llama31-saferlhf-stage3-os-s42 is an 8 billion parameter research checkpoint based on Meta's Llama-3.1-8B-Instruct model, developed using the RONPO OS Stage-3 method. This model is specifically intended for reproducibility and evaluation within the context of the RONPO paper, having passed a corrected stability gate and ranking first on a fixed panel by normalized worst-objective point estimate. It is not designed for production assistant use but serves as a research artifact for experimental validation.
Loading preview...
ronpo-llama31-saferlhf-stage3-os-s42: Research Checkpoint
This model, ronpo-llama31-saferlhf-stage3-os-s42, is an 8 billion parameter research checkpoint derived from the meta-llama/Llama-3.1-8B-Instruct base model. It was developed using the RONPO OS Stage-3 method as part of the RONPO AAAI revision experiments.
Key Characteristics & Evaluation
- Base Model:
meta-llama/Llama-3.1-8B-Instruct - Methodology: RONPO OS Stage-3
- Evaluation: Passed a corrected stability gate during Stage-3 SafeRLHF 1,000-prompt held-out evaluation.
- Performance: Ranked first on a fixed panel by the preregistered normalized worst-objective point estimate (0.4330; 95% interval [0.4151, 0.4508]). Its paired difference versus IPO was 0.0058 with 95% CI [-0.0183, 0.0295], indicating no statistically significant lead over IPO.
Intended Use
This checkpoint is primarily intended for reproducibility and evaluation related to the RONPO paper. It serves as a specific research artifact for experimental validation and is not designed for use as a production assistant.