princeton-nlp/Llama-3-Instruct-8B-SimPO-v0.2

Hugging Face
TEXT GENERATIONPricing:Input $0.37 / Cached $0.074 / Output $0.38Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:8kTool Calling:SupportedPublished:Jul 6, 2024Architecture:Transformer0.0K Featherless Exclusive Warm

Llama-3-Instruct-8B-SimPO-v0.2 is an 8 billion parameter instruction-tuned language model developed by princeton-nlp, based on the Llama 3 architecture. This model is fine-tuned using the SimPO (Simple Preference Optimization) method, which utilizes a reference-free reward mechanism. It is designed for general instruction following tasks, leveraging its 8192 token context length for processing longer inputs.

Loading preview...

Overview

Llama-3-Instruct-8B-SimPO-v0.2 is an 8 billion parameter instruction-tuned model from princeton-nlp, built upon the Llama 3 architecture. Its key differentiator is the application of SimPO (Simple Preference Optimization), a novel fine-tuning approach detailed in the preprint. SimPO distinguishes itself by employing a reference-free reward mechanism, which simplifies the preference optimization process.

Key Capabilities

  • Instruction Following: Designed to accurately follow user instructions for various tasks.
  • SimPO Fine-tuning: Leverages a unique preference optimization method that does not require reference responses.
  • Llama 3 Base: Benefits from the strong foundational capabilities of the Llama 3 8B model.
  • Extended Context: Supports an 8192 token context window, enabling processing of longer prompts and conversations.

Good For

  • Developers interested in exploring models fine-tuned with novel preference optimization techniques.
  • General-purpose instruction-following applications where a robust 8B parameter model is suitable.
  • Research into efficient and reference-free alignment methods for large language models.

Popular Sampler Settings

Top 3 parameter combinations used by Featherless users for this model. Click a tab to see each config.

temperature
top_p
top_k
frequency_penalty
presence_penalty
repetition_penalty
min_p