promotion/Llama-3.1-8B-NBPO-UltraFeedback-eta0.05
promotion/Llama-3.1-8B-NBPO-UltraFeedback-eta0.05 is an 8 billion parameter language model based on the Llama-3.1-8B-Instruct backbone, fine-tuned using Nash Bargaining Preference Optimization (NBPO) on the UltraFeedback dataset. This model is specifically optimized for instruction following, truthfulness, honesty, and helpfulness, utilizing a unique eta_t value of 0.05. It is designed for applications requiring highly aligned and reliable conversational AI outputs.
Loading preview...
Model Overview
promotion/Llama-3.1-8B-NBPO-UltraFeedback-eta0.05 is an 8 billion parameter language model built upon the meta-llama/Llama-3.1-8B-Instruct backbone. It has been fine-tuned using the Nash Bargaining Preference Optimization (NBPO) method, specifically targeting the UltraFeedback dataset. This model distinguishes itself by employing a unique eta_t value of 0.05, which differs from the released eta_t = 1 for its panel.
Key Capabilities
- Optimized Objectives: Focuses on improving instruction following, truthfulness, honesty, and helpfulness.
- Training Methodology: Utilizes Nash Bargaining Preference Optimization (NBPO) with a training budget of 300 optimizer updates and a global batch size of 16.
- Evaluation: Performance is assessed by independent objective-wise win rates against a common reference, judged by
Llama-3.3-70B-Instructon prompt-disjoint held-out prompts.
Good For
- Applications requiring a highly aligned and reliable conversational agent.
- Use cases where strong instruction following and truthful responses are critical.
- Scenarios benefiting from a model fine-tuned with a specific
eta_tvalue for enhanced preference optimization.