promotion/Llama-3.1-8B-SafeRLHF-Utilitarian-baseline
The promotion/Llama-3.1-8B-SafeRLHF-Utilitarian-baseline model is an 8 billion parameter language model based on Meta's Llama-3.1-8B-Instruct architecture. It is fine-tuned using a Utilitarian objective on the SafeRLHF panel, focusing on balancing helpfulness and harmlessness. This model is designed for applications requiring responses that prioritize both utility and safety, offering a distinct approach to objective weighting compared to other RLHF methods. It is particularly suited for scenarios where a balanced, ethical output is critical.
Loading preview...
Utilitarian on SafeRLHF
This model, promotion/Llama-3.1-8B-SafeRLHF-Utilitarian-baseline, is an 8 billion parameter language model built upon the meta-llama/Llama-3.1-8B-Instruct backbone. It distinguishes itself through its training methodology, employing a Utilitarian objective within the SafeRLHF framework. This approach focuses on achieving a balanced optimization of both helpfulness and harmlessness as primary objectives.
Key Characteristics
- Objective-wise Control: It utilizes a matched game-aggregation control, ensuring the same response pool, temperatures, disagreement point, and optimizer budget as NBPO, with the key difference being the utilitarian choice of objective weights.
- Training Details: The model underwent 300 optimizer updates with a global batch size of 16.
- Evaluation: Performance is assessed by an independent objective-wise win rate against a common reference policy. Evaluation is conducted by
Llama-3.3-70B-Instructon prompt-disjoint held-out prompts, considering both presentation orders.
Use Cases
This model is particularly relevant for applications where generating responses that are both highly useful and strictly adhere to safety guidelines is paramount. Its utilitarian alignment aims to provide a robust baseline for scenarios demanding a careful balance between performance and ethical considerations.