promotion/Llama-3.1-8B-NBPO-UltraFeedback-600updates
Llama-3.1-8B-NBPO-UltraFeedback-600updates is an 8 billion parameter language model developed by promotion, based on the Llama-3.1-8B-Instruct backbone. This model is fine-tuned using Nash Bargaining Preference Optimization (NBPO) on the UltraFeedback dataset, specifically optimized for instruction following, truthfulness, honesty, and helpfulness. It features a 32768 token context length and is designed for applications requiring robust adherence to instructions and ethical AI principles.
Loading preview...
Model Overview
This model, Llama-3.1-8B-NBPO-UltraFeedback-600updates, is an 8 billion parameter language model built upon the meta-llama/Llama-3.1-8B-Instruct backbone. It has been fine-tuned using the Nash Bargaining Preference Optimization (NBPO) method, specifically targeting the UltraFeedback dataset. The training involved 300 optimizer updates with a global batch size of 16, as detailed in the Appendix of the Nash Bargaining Preference Optimization paper regarding optimizer-budget sensitivity.
Key Capabilities
- Instruction Following: Optimized to accurately follow user instructions.
- Truthfulness & Honesty: Enhanced for generating factually correct and honest responses.
- Helpfulness: Designed to provide useful and relevant information.
- Robustness: Evaluated for objective-wise win rate against a common reference using
Llama-3.3-70B-Instructon prompt-disjoint held-out prompts.
Good For
- Applications requiring high fidelity in instruction adherence.
- Use cases where truthfulness and honesty are critical, such as factual Q&A or content generation.
- Scenarios demanding helpful and ethical AI interactions.