princeton-nlp/Llama-3-Instruct-8B-SimPO-v0.2
Llama-3-Instruct-8B-SimPO-v0.2 is an 8 billion parameter instruction-tuned language model developed by princeton-nlp, based on the Llama 3 architecture. This model is fine-tuned using the SimPO (Simple Preference Optimization) method, which utilizes a reference-free reward mechanism. It is designed for general instruction following tasks, leveraging its 8192 token context length for processing longer inputs.
Loading preview...
Overview
Llama-3-Instruct-8B-SimPO-v0.2 is an 8 billion parameter instruction-tuned model from princeton-nlp, built upon the Llama 3 architecture. Its key differentiator is the application of SimPO (Simple Preference Optimization), a novel fine-tuning approach detailed in the preprint. SimPO distinguishes itself by employing a reference-free reward mechanism, which simplifies the preference optimization process.
Key Capabilities
- Instruction Following: Designed to accurately follow user instructions for various tasks.
- SimPO Fine-tuning: Leverages a unique preference optimization method that does not require reference responses.
- Llama 3 Base: Benefits from the strong foundational capabilities of the Llama 3 8B model.
- Extended Context: Supports an 8192 token context window, enabling processing of longer prompts and conversations.
Good For
- Developers interested in exploring models fine-tuned with novel preference optimization techniques.
- General-purpose instruction-following applications where a robust 8B parameter model is suitable.
- Research into efficient and reference-free alignment methods for large language models.
Top 3 parameter combinations used by Featherless users for this model. Click a tab to see each config.