MaLA-LM/emma-500-llama3-8b-bi

Hugging Face
TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:8kTool Calling:SupportedPublished:May 10, 2025License:llama3Architecture:Transformer Featherless Exclusive Warm

MaLA-LM/emma-500-llama3-8b-bi is an 8 billion parameter multilingual language model developed by MaLA-LM, built upon the Llama 3 architecture. It is continually pre-trained on 671 billion tokens from the MaLA Corpus, supporting 546 languages with substantial data and over 2,500 bilingual translation pairs. This model excels in massively multilingual NLP tasks like machine translation and commonsense reasoning, particularly for low-resource languages.

Loading preview...

EMMA-500 Llama 3 8B Bilingual Model

EMMA-500 Llama 3 8B is a multilingual language model from MaLA-LM, continually pre-trained on the Llama 3 8B architecture. Its primary focus is to enhance language representation, especially for low-resource languages, by leveraging a diverse and massive dataset.

Key Capabilities & Features

  • Massively Multilingual: Supports 546 languages with over 100k tokens each, and includes bilingual translation data for over 2,500 language pairs.
  • Continual Pre-training: Built on Llama 3 8B, it has undergone extensive continual pre-training on 671 billion tokens.
  • Diverse Data Mix: Trained on the MaLA Corpus, which includes books, code, instruction data, and papers, with a specific focus on bilingual text.
  • Performance: Excels in multilingual tasks such as commonsense reasoning, machine translation, and text classification.

Use Cases

This model is particularly well-suited for:

  • Massively multilingual NLP tasks, including machine translation.
  • Research and development in low-resource language processing.

Limitations

  • May show performance regression on some tasks and high-resource languages.
  • Not recommended for real-world scenarios, especially in high-stakes domains.