Magpie-Align/Llama-3.1-8B-Magpie-Align-v0.2
Magpie-Align/Llama-3.1-8B-Magpie-Align-v0.2 is an 8 billion parameter aligned version of Meta's Llama-3.1-8B, developed by Magpie-Align. This model is fine-tuned using a two-stage pipeline involving supervised fine-tuning (SFT) with custom datasets and Direct Preference Optimization (DPO). It demonstrates enhanced alignment and performance compared to the official Llama-3.1-8B-Instruct, excelling in benchmarks like Alpaca Eval 2 and Arena Hard.
Loading preview...
Magpie-Align/Llama-3.1-8B-Magpie-Align-v0.2 Overview
This model is an aligned variant of the 8 billion parameter Meta-Llama-3.1-8B developed by Magpie-Align. It leverages a unique alignment pipeline designed to improve performance over the base Llama 3.1 instruction model.
Key Capabilities and Training
The model's alignment process involves two main stages:
- Supervised Fine-tuning (SFT): Initial SFT was performed using proprietary datasets, specifically Magpie-Align/Magpie-Llama-3.1-Pro-500K-Filtered and Magpie-Align/Magpie-Reasoning-150K. The SFT checkpoint is available as Magpie-Align/Llama-3.1-8B-Magpie-Align-SFT-v0.2.
- Direct Preference Optimization (DPO): Following SFT, DPO was applied using the Magpie-Align/Llama-3.1-70B-PO-100K-armorm dataset, which is based on the ArmoRM reward model.
Performance Highlights
This model demonstrates improved performance compared to the official Llama-3.1-8B-Instruct, with notable results:
- Alpaca Eval 2 (vs GPT-4-Turbo-1106): 46.68 (LC), 53.42 (WR)
- Arena Hard: 43.2
Usage and Licensing
Users should employ the Llama 3 official chat template for optimal performance. The model adheres to the Meta Llama 3.1 Community License. For detailed instructions on how to use the model, refer to the official Llama 3.1 repository and replace the model ID with this one.