promotion/Llama-3.1-8B-SafeRLHF-NBPO-finitepool
The promotion/Llama-3.1-8B-SafeRLHF-NBPO-finitepool model is an 8 billion parameter language model based on meta-llama/Llama-3.1-8B-Instruct, fine-tuned using Nash Bargaining Preference Optimization (NBPO) with a finite-pool realization. It is specifically optimized for balancing helpfulness and harmlessness, building upon the SafeRLHF panel. This model is designed for applications requiring robust safety alignment alongside strong instructional capabilities, featuring a 32,768 token context length.
Loading preview...
Overview
The promotion/Llama-3.1-8B-SafeRLHF-NBPO-finitepool model is an 8 billion parameter language model derived from meta-llama/Llama-3.1-8B-Instruct. It has been fine-tuned using a specific implementation of Nash Bargaining Preference Optimization (NBPO), referred to as the finite-pool realization of Algorithm 1. This optimization process, conducted with 300 optimizer updates and a global batch of 16, focuses on achieving a balance between helpfulness and harmlessness within the SafeRLHF framework.
Key Capabilities
- Safety Alignment: Optimized for both helpfulness and harmlessness objectives, making it suitable for sensitive applications.
- Preference Optimization: Utilizes Nash Bargaining Preference Optimization (NBPO) for fine-tuning, a method detailed in its corresponding research.
- Robust Evaluation: Performance is assessed through independent objective-wise win rates against a common reference, judged by
Llama-3.3-70B-Instructon held-out prompts.
Good for
- Applications requiring a strong emphasis on safety and ethical AI behavior.
- Use cases where balancing helpfulness with harmlessness is critical.
- Developers seeking a model fine-tuned with advanced preference optimization techniques for aligned responses.