MaLA-LM/emma-500-llama3.1-8b-mono

TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:8kTool Calling:SupportedPublished:May 10, 2025License:llama3Architecture:Transformer Featherless Exclusive Cold

MaLA-LM/emma-500-llama3.1-8b-mono is an 8 billion parameter multilingual language model developed by MaLA-LM, continually pre-trained on the Llama 3.1 architecture. It supports 546 languages with substantial training data, leveraging the MaLA Corpus which includes books, code, instruction data, and papers. This model excels in multilingual tasks such as commonsense reasoning, machine translation, and text classification, particularly for low-resource languages.

Loading preview...

EMMA-500 Llama 3.1 8B: Massively Multilingual Adaptation

EMMA-500 Llama 3.1 8B is a multilingual language model from MaLA-LM, built upon the Llama 3.1 8B architecture. It has undergone continual pre-training to enhance language representation, especially for low-resource languages, using a diverse dataset of 419 billion tokens.

Key Capabilities and Features

  • Massive Multilingual Support: Supports 546 languages, each with over 100k tokens of training data, making it highly effective for diverse linguistic tasks.
  • Continual Pre-training: Leverages the extensive MaLA Corpus, which includes a rich mix of monolingual data from domains like code, books, instruction data, and academic papers.
  • Optimized for Multilingual Tasks: Designed to excel in tasks such as commonsense reasoning, machine translation, and text classification across many languages.
  • Base Model: Built on the robust Llama 3.1 8B foundation, enhanced for multilingual performance.

Use Cases and Considerations

This model is particularly suited for:

  • Massively multilingual NLP tasks, including machine translation.
  • Research and development in low-resource language processing.

It is important to note that the model may exhibit performance regression on some tasks and high-resource languages compared to models specifically trained for those, and it is not recommended for real-world, high-stakes scenarios.