promotion/Llama-3.1-8B-NBPO-UltraFeedback-clean
promotion/Llama-3.1-8B-NBPO-UltraFeedback-clean is an 8 billion parameter language model based on the Llama-3.1-8B-Instruct backbone, developed by promotion. It is fine-tuned using Nash Bargaining Preference Optimization (NBPO) on a contamination-free UltraFeedback dataset, specifically optimized for instruction following, truthfulness, honesty, and helpfulness. This model is designed for applications requiring robust and reliable responses across various objective-wise criteria.
Loading preview...
Model Overview
This model, promotion/Llama-3.1-8B-NBPO-UltraFeedback-clean, is an 8 billion parameter language model built upon the meta-llama/Llama-3.1-8B-Instruct backbone. It has been fine-tuned using the Nash Bargaining Preference Optimization (NBPO) method, as detailed in the NBPO paper (Table 2).
Key Capabilities
- Instruction Following: Optimized to accurately follow user instructions.
- Truthfulness & Honesty: Enhanced for generating factually correct and honest responses.
- Helpfulness: Designed to provide useful and relevant information.
- Contamination-Free Training: Retrained on a filtered UltraFeedback dataset where pairs are prompt-disjoint from evaluation prompts, ensuring robust and unbiased performance.
Training Details
The model underwent 300 optimizer updates with a global batch size of 16, using an Adafactor optimizer and a learning rate of 5e-7. The training specifically targeted the UltraFeedback panel, focusing on the aforementioned objectives.
Evaluation
Evaluation was conducted using an independent objective-wise win rate against a common reference model. Responses were judged by Llama-3.3-70B-Instruct on held-out, prompt-disjoint prompts, considering both presentation orders to ensure fairness.