promotion/Llama-3.1-8B-NBPO-finitepool

TEXT GENERATIONPricing:Input $0.2 / Cached $0.028 / Output $0.32Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 5, 2026License:llama3.1Architecture:Transformer Featherless Exclusive Cold

promotion/Llama-3.1-8B-NBPO-finitepool is an 8 billion parameter language model with a 32768 token context length. This model was trained using NBPO on UltraFeedback, but it contains a contaminated training split where evaluation prompts were included in the training data. It is superseded by a clean version and is provided for inspectability rather than general use.

Loading preview...

Model Overview

promotion/Llama-3.1-8B-NBPO-finitepool is an 8 billion parameter language model with a 32768 token context length, developed by promotion. It was trained using the NBPO (Neural Best-of-N Policy Optimization) method on the UltraFeedback dataset.

Important Note: Contaminated Training Data

This specific checkpoint is superseded and should not be used for general applications or for reproducing the UltraFeedback results presented in the NBPO paper. An audit revealed that 629 of the 698 UltraFeedback evaluation prompts were inadvertently included in this checkpoint's training pairs. This contamination means the model was partly evaluated on data it had already seen, leading to a biased performance measurement.

Replacement Model

The correct and prompt-disjoint version, promotion/Llama-3.1-8B-NBPO-UltraFeedback-clean, was retrained with a filtered dataset (34,026 training rows over 5,671 prompts) to ensure no overlap with the evaluation set. This clean version is the one referenced in the NBPO paper for UltraFeedback results.

Intended Use

This promotion/Llama-3.1-8B-NBPO-finitepool checkpoint is retained solely for inspectability as a published artifact, allowing researchers to examine the original, contaminated training run. It is not recommended for deployment or performance evaluation.