Magpie-Align/Llama-3.1-8B-Magpie-Align-v0.1
Magpie-Align/Llama-3.1-8B-Magpie-Align-v0.1 is an 8 billion parameter language model developed by Zhangchen Xu et al., built upon the Llama-3.1 architecture. This model is aligned using a two-stage process involving Supervised Fine-tuning (SFT) and Direct Preference Optimization (DPO), leveraging a self-synthesis method called Magpie for high-quality instruction data generation. It excels in alignment benchmarks like AlpacaEval, ArenaHard, and WildBench, making it suitable for applications requiring robust instruction following and preference alignment.
Loading preview...
Magpie-Align/Llama-3.1-8B-Magpie-Align-v0.1 Overview
This model is an 8 billion parameter language model based on the Llama-3.1 architecture, developed by Zhangchen Xu et al. It undergoes a two-stage alignment pipeline: Supervised Fine-tuning (SFT) and Direct Preference Optimization (DPO). A key innovation is the use of the Magpie self-synthesis method for generating high-quality instruction data, which involves prompting aligned LLMs like Llama-3-Instruct to create user queries and responses at scale.
Key Capabilities & Training
- Data Synthesis: Utilizes the Magpie method to synthesize 4 million instructions and responses, from which 300K high-quality instances are selected for training.
- Alignment Performance: Models fine-tuned with Magpie data for SFT have shown to surpass previous public datasets used for both SFT and preference optimization (e.g., UltraFeedback) on alignment benchmarks.
- Benchmark Results: Demonstrates strong performance on alignment benchmarks such as AlpacaEval, ArenaHard, and WildBench.
- Training Process: Employs DPO with specific hyperparameters including a learning rate of 1e-06, a total batch size of 256, and 1 epoch of training.
When to Use This Model
- Instruction Following: Ideal for applications requiring robust and nuanced instruction adherence.
- Preference Alignment: Suitable for tasks where aligning with human preferences is critical.
- Research on Alignment: A valuable resource for researchers exploring data synthesis and alignment techniques, particularly the Magpie method.