Magpie-Align/Llama-3.1-8B-Magpie-Align-v0.1

TEXT GENERATIONPricing:Input $0.2 / Cached $0.028 / Output $0.32Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 24, 2024License:llama3.1Architecture:Transformer0.0K Featherless Exclusive Cold

Magpie-Align/Llama-3.1-8B-Magpie-Align-v0.1 is an 8 billion parameter language model developed by Zhangchen Xu et al., built upon the Llama-3.1 architecture. This model is aligned using a two-stage process involving Supervised Fine-tuning (SFT) and Direct Preference Optimization (DPO), leveraging a self-synthesis method called Magpie for high-quality instruction data generation. It excels in alignment benchmarks like AlpacaEval, ArenaHard, and WildBench, making it suitable for applications requiring robust instruction following and preference alignment.

Loading preview...

Magpie-Align/Llama-3.1-8B-Magpie-Align-v0.1 Overview

This model is an 8 billion parameter language model based on the Llama-3.1 architecture, developed by Zhangchen Xu et al. It undergoes a two-stage alignment pipeline: Supervised Fine-tuning (SFT) and Direct Preference Optimization (DPO). A key innovation is the use of the Magpie self-synthesis method for generating high-quality instruction data, which involves prompting aligned LLMs like Llama-3-Instruct to create user queries and responses at scale.

Key Capabilities & Training

  • Data Synthesis: Utilizes the Magpie method to synthesize 4 million instructions and responses, from which 300K high-quality instances are selected for training.
  • Alignment Performance: Models fine-tuned with Magpie data for SFT have shown to surpass previous public datasets used for both SFT and preference optimization (e.g., UltraFeedback) on alignment benchmarks.
  • Benchmark Results: Demonstrates strong performance on alignment benchmarks such as AlpacaEval, ArenaHard, and WildBench.
  • Training Process: Employs DPO with specific hyperparameters including a learning rate of 1e-06, a total batch size of 256, and 1 epoch of training.

When to Use This Model

  • Instruction Following: Ideal for applications requiring robust and nuanced instruction adherence.
  • Preference Alignment: Suitable for tasks where aligning with human preferences is critical.
  • Research on Alignment: A valuable resource for researchers exploring data synthesis and alignment techniques, particularly the Magpie method.