promotion/Llama-3.1-8B-SafeRLHF-SurplusMaxMin-baseline
The promotion/Llama-3.1-8B-SafeRLHF-SurplusMaxMin-baseline model is an 8 billion parameter language model based on the Llama-3.1-8B-Instruct backbone, fine-tuned using the SafeRLHF framework with a SurplusMaxMin objective. It is designed to balance helpfulness and harmlessness, utilizing a 32768 token context length. This model is optimized for controlled response generation, aiming for robust performance in safety-critical applications.
Loading preview...
Overview
This model, promotion/Llama-3.1-8B-SafeRLHF-SurplusMaxMin-baseline, is an 8 billion parameter language model built upon the meta-llama/Llama-3.1-8B-Instruct backbone. It has been fine-tuned using the SafeRLHF framework, specifically employing a SurplusMaxMin objective. The training focused on balancing helpfulness and harmlessness, with a budget of 300 optimizer updates and a global batch size of 16.
Key Capabilities
- Safety-Oriented Fine-tuning: Utilizes the SafeRLHF panel with SurplusMaxMin for objective weighting, aiming to control response generation for safety.
- Objective Balancing: Designed to optimize for both helpfulness and harmlessness simultaneously.
- Controlled Training: Matched game-aggregation control during training, ensuring a consistent response pool, temperatures, disagreement points, and optimizer budget compared to NBPO.
Evaluation
The model's performance is evaluated through independent objective-wise win rates against a common reference. This assessment is conducted by Llama-3.3-70B-Instruct on prompt-disjoint held-out prompts, considering both presentation orders. This evaluation protocol is detailed in the Nash Bargaining Preference Optimization (NBPO) paper, Table 2, as part of its primary cross-method evaluation.