princeton-nlp/Llama-3-Base-8B-SFT-DPO
TEXT GENERATIONConcurrency Cost:1Model Size:8BQuant:FP8Ctx Length:8kPublished:May 17, 2024Architecture:Transformer Warm
The princeton-nlp/Llama-3-Base-8B-SFT-DPO is an 8 billion parameter Llama-3-based language model developed by Princeton NLP, fine-tuned using the SimPO (Simple Preference Optimization with a Reference-Free Reward) method. This model is specifically optimized for preference alignment without requiring a reference reward model, making it suitable for tasks benefiting from direct preference optimization. It offers an 8192-token context window and is derived from research detailed in the SimPO preprint.
Loading preview...
Popular Sampler Settings
Top 3 parameter combinations used by Featherless users for this model. Click a tab to see each config.
temperature
top_p
top_k
frequency_penalty
presence_penalty
repetition_penalty
min_p