promotion/Llama-3.1-8B-TLDR-NBPO-finitepool

TEXT GENERATIONPricing:Input $0.2 / Cached $0.028 / Output $0.32Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 5, 2026License:llama3.1Architecture:Transformer Featherless Exclusive Cold

Llama-3.1-8B-TLDR-NBPO-finitepool is an 8 billion parameter model based on Meta's Llama-3.1-8B-Instruct, fine-tuned using Nash Bargaining Preference Optimization (NBPO). This model is specifically optimized for summarization tasks, focusing on coverage, faithfulness, conciseness, and helpfulness. It is designed to generate high-quality summaries, making it suitable for applications requiring efficient information distillation from longer texts. The model leverages a finite-pool realization of the NBPO algorithm for its training methodology.

Loading preview...

Model Overview

promotion/Llama-3.1-8B-TLDR-NBPO-finitepool is an 8 billion parameter language model derived from meta-llama/Llama-3.1-8B-Instruct. It has been fine-tuned using a specific method called Nash Bargaining Preference Optimization (NBPO), specifically its finite-pool realization. This model is tailored for summarization tasks, particularly on the TL;DR panel.

Key Capabilities and Objectives

The primary objectives guiding this model's development and performance are:

  • Coverage: Ensuring all important aspects of the source text are included.
  • Faithfulness: Maintaining accuracy and not introducing new information.
  • Conciseness: Producing brief and to-the-point summaries.
  • Helpfulness: Generating summaries that are useful and informative to the user.

Training and Evaluation

The model underwent 300 optimizer updates with a global batch size of 16. Its performance is evaluated based on an independent objective-wise win rate against the common reference policy (Llama-3.1-8B-Instruct). Evaluation is conducted by Llama-3.3-70B-Instruct on prompt-disjoint held-out prompts, considering both presentation orders. This model is documented in the "Nash Bargaining Preference Optimization (NBPO)" paper, specifically in Table 2, which details its cross-method evaluation.