togethercomputer/gemma-2-9b-it-MoAA-DPO
togethercomputer/gemma-2-9b-it-MoAA-DPO is a 9 billion parameter instruction-tuned causal language model, fine-tuned from Gemma-2-9b-it using the Mixture of Agents Alignment (MoAA) pipeline. Developed by togethercomputer, this model leverages collective intelligence from open-source LLMs for alignment, demonstrating significant improvements on benchmarks like Arena-Hard. It is primarily designed for advanced alignment tasks and self-improvement in LLM performance.
Loading preview...
Model Overview
This model, togethercomputer/gemma-2-9b-it-MoAA-DPO, is a 9 billion parameter instruction-tuned variant of Gemma-2-9b-it. It is specifically fine-tuned using the Mixture of Agents Alignment (MoAA) pipeline, an approach developed by togethercomputer that utilizes the collective intelligence of open-source LLMs to enhance alignment.
Key Features and Alignment Method
MoAA involves two main stages:
- Synthetic Data Generation: Employs Mixture of Agents (MoA) to produce high-quality synthetic data for supervised fine-tuning.
- Preference Annotation: Combines multiple LLMs as a reward model to provide preference annotations for DPO training.
Performance and Impact
The MoAA method has shown substantial improvements in alignment:
- Benchmark Gains: Increased Llama-3.1-8B-Instruct's Arena-Hard score from 19 to 48, and Gemma-2-9B-it's score from 42 to 56, outperforming GPT-4o-labeled sets at the time.
- Ensembled Rewards: An MoA reward model with dynamic criteria filtering proved more effective than competitive ArmoRM on MT-Bench and Arena-Hard, while remaining entirely open source.
- Self-Improvement: Models fine-tuned on MoAA data have demonstrated the ability to surpass their own teachers, indicating a pathway for open models to exceed proprietary performance without external supervision.
Use Cases
This model is particularly well-suited for research and applications focused on:
- Advanced LLM alignment techniques.
- Leveraging collective intelligence for model improvement.
- Developing self-improving AI systems.
For detailed evaluation metrics and further information, refer to the accompanying paper.