promotion/Llama-3.1-8B-TLDR-NBPO-eta0.05
Llama-3.1-8B-TLDR-NBPO-eta0.05 is an 8 billion parameter Llama-3.1-8B-Instruct backbone model, fine-tuned using Nash Bargaining Preference Optimization (NBPO) with an eta_t of 0.05. This model is specifically optimized for generating concise, faithful, and helpful summaries, focusing on coverage and conciseness. It is designed for tasks requiring high-quality summarization from longer texts.
Loading preview...
Model Overview
This model, Llama-3.1-8B-TLDR-NBPO-eta0.05, is an 8 billion parameter language model built upon the meta-llama/Llama-3.1-8B-Instruct backbone. It has been fine-tuned using the Nash Bargaining Preference Optimization (NBPO) method, specifically with an eta_t value of 0.05. This particular eta_t differs from the initially released value of 1, indicating a targeted optimization approach.
Key Capabilities
The model is designed to excel in summarization tasks, focusing on the following objectives:
- Coverage: Ensuring all important aspects of the source material are included.
- Faithfulness: Maintaining accuracy and consistency with the original text.
- Conciseness: Producing brief and to-the-point summaries.
- Helpfulness: Generating summaries that are useful and informative to the user.
Training and Evaluation
- Training Budget: The model underwent 300 optimizer updates with a global batch size of 16.
- Evaluation: Performance is assessed through independent objective-wise win rates against a common reference model, judged by
Llama-3.3-70B-Instructon prompt-disjoint held-out prompts. This evaluation considers both presentation orders to ensure robust assessment.
Good For
- Generating high-quality, concise summaries of documents or conversations.
- Applications requiring faithful and helpful text condensation.
- Use cases where balancing coverage and brevity is crucial.