promotion/Llama-3.1-8B-NBPO-UltraFeedback-clean

TEXT GENERATIONPricing:Input $0.2 / Cached $0.028 / Output $0.32Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 6, 2026License:llama3.1Architecture:Transformer Featherless Exclusive Cold

promotion/Llama-3.1-8B-NBPO-UltraFeedback-clean is an 8 billion parameter language model based on the Llama-3.1-8B-Instruct backbone, developed by promotion. It is fine-tuned using Nash Bargaining Preference Optimization (NBPO) on a contamination-free UltraFeedback dataset, specifically optimized for instruction following, truthfulness, honesty, and helpfulness. This model is designed for applications requiring robust and reliable responses across various objective-wise criteria.

Loading preview...

Model Overview

This model, promotion/Llama-3.1-8B-NBPO-UltraFeedback-clean, is an 8 billion parameter language model built upon the meta-llama/Llama-3.1-8B-Instruct backbone. It has been fine-tuned using the Nash Bargaining Preference Optimization (NBPO) method, as detailed in the NBPO paper (Table 2).

Key Capabilities

  • Instruction Following: Optimized to accurately follow user instructions.
  • Truthfulness & Honesty: Enhanced for generating factually correct and honest responses.
  • Helpfulness: Designed to provide useful and relevant information.
  • Contamination-Free Training: Retrained on a filtered UltraFeedback dataset where pairs are prompt-disjoint from evaluation prompts, ensuring robust and unbiased performance.

Training Details

The model underwent 300 optimizer updates with a global batch size of 16, using an Adafactor optimizer and a learning rate of 5e-7. The training specifically targeted the UltraFeedback panel, focusing on the aforementioned objectives.

Evaluation

Evaluation was conducted using an independent objective-wise win rate against a common reference model. Responses were judged by Llama-3.3-70B-Instruct on held-out, prompt-disjoint prompts, considering both presentation orders to ensure fairness.