togethercomputer/Llama-3.1-8B-Instruct-MoAA-SFT

TEXT GENERATIONPricing:Input $0.2 / Cached $0.028 / Output $0.32Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 20, 2024Architecture:Transformer Featherless Exclusive Cold

togethercomputer/Llama-3.1-8B-Instruct-MoAA-SFT is an 8 billion parameter instruction-tuned model, based on Llama-3.1-8B-Instruct, developed by togethercomputer. It leverages a Mixture of Agents Alignment (MoAA) pipeline for supervised fine-tuning, utilizing collective intelligence from open-source LLMs to generate high-quality synthetic data. This model demonstrates significant alignment improvements, outperforming GPT-4o-labeled sets on benchmarks like Arena-Hard, and is designed for advanced alignment tasks.

Loading preview...

Model Overview

togethercomputer/Llama-3.1-8B-Instruct-MoAA-SFT is an 8 billion parameter instruction-tuned model built upon Llama-3.1-8B-Instruct. It is the supervised fine-tuning (SFT) component of the Mixture of Agents Alignment (MoAA) pipeline, a novel approach developed by togethercomputer that enhances model alignment by leveraging the collective intelligence of multiple open-source LLMs.

Key Capabilities & Innovations

  • Advanced Alignment: The MoAA method significantly improves alignment, boosting Llama-3.1-8B-Instruct's Arena-Hard score from 19 to 48 and Gemma-2-9B-it from 42 to 56, surpassing results from GPT-4o-labeled datasets.
  • Ensembled Reward Models: Utilizes a Mixture of Agents (MoA) reward model with dynamic criteria filtering, which outperforms competitive ArmoRM on MT-Bench and Arena-Hard, while remaining entirely open-source.
  • Self-Improvement: The model demonstrates self-improvement, as fine-tuning the strongest ensemble member on MoAA-generated data allows it to exceed the performance of its "teachers," indicating a path for open models to advance beyond proprietary ceilings without external supervision.
  • Synthetic Data Generation: Employs MoA to produce high-quality synthetic data for SFT, drawing from datasets like UltraFeedback and UltraChat, with proposers including WizardLM-2-8x22b and Llama-3.1-70b-Instruct, and Qwen-1.5-110b-Instruct as the aggregator.

Use Cases

This model is particularly well-suited for applications requiring highly aligned and robust instruction following, especially in scenarios where leveraging collective intelligence for improved performance is beneficial. Developers can use it for tasks demanding strong performance on benchmarks like Arena-Hard and MT-Bench, or for exploring advanced alignment techniques.