promotion/Llama-3.1-8B-SafeRLHF-NBPO-600updates
This model is a Llama-3.1-8B-Instruct backbone fine-tuned with 8 billion parameters and a 32768 token context length, developed for SafeRLHF using Nash Bargaining Preference Optimization (NBPO). It is specifically optimized for balancing helpfulness and harmlessness, demonstrating improved performance in win rate and held-out fit compared to other budget allocations. The model is designed for applications requiring robust safety and balanced objective optimization.
Loading preview...
Model Overview
This model, promotion/Llama-3.1-8B-SafeRLHF-NBPO-600updates, is built upon the meta-llama/Llama-3.1-8B-Instruct backbone, featuring 8 billion parameters and a 32768 token context length. It has been fine-tuned using Nash Bargaining Preference Optimization (NBPO) within a SafeRLHF framework, specifically undergoing 300 optimizer updates with a global batch size of 16.
Key Capabilities
- Objective Balancing: Optimized for simultaneously achieving high helpfulness and harmlessness, as detailed in the Nash Bargaining Preference Optimization (NBPO) appendix.
- Performance: Demonstrates superior held-out fit and win rate compared to models trained with different optimization budgets, with an nMSE of 0.9591 and a worst-case win rate of 0.696.
- Safety-Oriented Fine-tuning: Leverages SafeRLHF to ensure robust safety characteristics alongside utility.
Good For
- Safety-Critical Applications: Ideal for use cases where balancing helpfulness with strict harmlessness is paramount.
- Research in RLHF: Provides a strong baseline for further research into preference optimization and safety alignment in large language models.
- Balanced AI Assistants: Suitable for developing conversational agents that require careful consideration of both utility and ethical guidelines.