MaLA-LM/emma-500-llama2-7b

Hugging Face
TEXT GENERATIONConcurrent Unit Cost:1Model Size:7BQuant:FP8Context Size:4kPublished:Sep 14, 2024License:llama2Architecture:Transformer0.0K Open Weights Featherless Exclusive Warm

EMMA-500 is a multilingual language model developed by MaLA-LM, built on the Llama 2 7B architecture. It is continually pre-trained on the MaLA Corpus, encompassing over 500 languages and 74 billion tokens, to enhance language representation, especially in low-resource languages. This model excels in diverse multilingual tasks such as commonsense reasoning, machine translation, text classification, and open-ended generation, outperforming other Llama 2-based models in these areas.

Loading preview...

EMMA-500: Enhanced Multilingual Adaptation

EMMA-500 is a multilingual language model developed by MaLA-LM, based on the Llama 2 7B architecture. It significantly improves language representation, particularly for low-resource languages, through continual pre-training on the extensive MaLA Corpus. This corpus comprises over 74 billion tokens across more than 500 languages.

Key Capabilities & Performance

  • Massively Multilingual: Supports 546 languages, each with substantial training data (over 100k tokens).
  • Diverse Task Proficiency: Excels in a wide range of tasks including commonsense reasoning, machine translation, text classification, natural language inference, code generation, and open-ended generation.
  • Outperforms Llama 2-based models: Demonstrates superior performance in diverse multilingual settings, particularly in text classification and natural language inference.
  • Intrinsic Evaluation: Achieves the lowest negative log-likelihood in intrinsic evaluations.
  • Enhanced Code Generation: Shows improved performance in code generation and machine reading comprehension (MRC).

Training & Data

EMMA-500 leverages a diverse data mix including code, books, and instruction data. The model's development involved researchers from Helsinki-NLP, TU Darmstadt, the University of Edinburgh, and LMU Munich, funded by HPLT and UTTER.

Considerations

Challenges remain in low-resource languages, where the model may exhibit higher Self-BLEU scores, indicating reduced output diversity.