MaLA-LM/emma-500-llama3.1-8b-bi

TEXT GENERATIONPricing:Input $0.37 / Cached $0.074 / Output $0.38Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:8kTool Calling:SupportedPublished:May 10, 2025License:llama3Architecture:Transformer0.0K Featherless Exclusive Cold

The MaLA-LM/emma-500-llama3.1-8b-bi is an 8 billion parameter multilingual language model developed by MaLA-LM, built upon the Llama 3.1 architecture. It is continually pre-trained on 671 billion tokens from the MaLA Corpus, supporting 546 languages with substantial data and over 2,500 bilingual translation pairs. This model excels in massively multilingual NLP tasks like machine translation and commonsense reasoning, particularly enhancing language representation in low-resource languages.

Loading preview...

EMMA-500 Llama 3.1 8B: Massively Multilingual Adaptation

EMMA-500 Llama 3.1 8B is an 8 billion parameter multilingual language model from MaLA-LM, continually pre-trained on the Llama 3.1 architecture. It leverages the extensive MaLA Corpus, which includes over 500 languages, augmented with books, code, instruction data, and papers.

Key Capabilities & Features

  • Massively Multilingual: Supports 546 languages with significant training data (over 100k tokens each).
  • Bilingual Translation Data: Uniquely trained with bilingual translation data across more than 2,500 language pairs, enhancing cross-lingual understanding.
  • Diverse Data Mix: Pre-trained on a 671 billion token dataset comprising monolingual text, code, and bilingual translation data.
  • Architecture: Built on the robust Llama 3.1 8B base model, continually adapted for improved language representation, especially in low-resource languages.

Use Cases

  • Multilingual NLP Tasks: Ideal for applications requiring understanding and generation across many languages, such as machine translation and commonsense reasoning.
  • Low-Resource Language Support: Particularly beneficial for improving performance in languages with limited existing data.

Note: The model may show performance regression on some tasks and high-resource languages and is not recommended for real-world, high-stakes scenarios.