togethercomputer/gemma-2-9b-it-MoAA-SFT

TEXT GENERATIONPricing:Input $0.431 / Output $1.12Concurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:16kPublished:Sep 23, 2024Architecture:Transformer0.0K Featherless Exclusive Cold

The togethercomputer/gemma-2-9b-it-MoAA-SFT model is a 9 billion parameter instruction-tuned variant of Gemma-2, developed by Together Computer. It is fine-tuned using the Mixture of Agents Alignment (MoAA) pipeline, which leverages collective intelligence from open-source LLMs to generate high-quality synthetic data for supervised fine-tuning. This model demonstrates significant alignment improvements, outperforming GPT-4o-labeled sets on benchmarks like Arena-Hard, and is particularly suited for tasks requiring advanced alignment and robust performance derived from ensemble-based training.

Loading preview...

Model Overview

This model, gemma-2-9b-it-MoAA-SFT, is a 9 billion parameter instruction-tuned version of Gemma-2, developed by Together Computer. It is the Supervised Fine-Tuning (SFT) component of their Mixture of Agents Alignment (MoAA) pipeline, an innovative approach that utilizes the collective intelligence of multiple open-source LLMs to enhance model alignment.

Key Capabilities and Innovations

  • Advanced Alignment: The MoAA method significantly improves alignment, with the model showing substantial gains on benchmarks like Arena-Hard (Gemma-2-9B-it improved from 42 to 56), surpassing results from GPT-4o-labeled datasets.
  • Ensembled Rewards: MoAA employs an ensemble of LLMs as a reward model, incorporating dynamic criteria filtering. This approach has been shown to outperform competitive ArmoRM on MT-Bench and Arena-Hard, while maintaining a fully open-source methodology.
  • Self-Improvement: The model demonstrates self-improvement, as fine-tuning the strongest ensemble member on MoAA-generated data allows it to surpass its "teachers," indicating a pathway for open models to exceed proprietary performance without external supervision.

Training Details

The model was fine-tuned using a subsampled dataset from UltraFeedback and UltraChat. The synthetic responses for training were generated using a Mixture of Agents (MoA) approach, where proposers included WizardLM-2-8x22b, Gemma-2-7b-it, Qwen-2-72b-Instruct, and Llama-3.1-70b-Instruct, with Qwen-1.5-110b-Instruct serving as the aggregator. For detailed evaluation metrics, refer to the accompanying paper.