MaLA-LM/emma-500-llama2-7b
EMMA-500 is a multilingual language model developed by MaLA-LM, built on the Llama 2 7B architecture. It is continually pre-trained on the MaLA Corpus, encompassing over 500 languages and 74 billion tokens, to enhance language representation, especially in low-resource languages. This model excels in diverse multilingual tasks such as commonsense reasoning, machine translation, text classification, and open-ended generation, outperforming other Llama 2-based models in these areas.
Loading preview...
EMMA-500: Enhanced Multilingual Adaptation
EMMA-500 is a multilingual language model developed by MaLA-LM, based on the Llama 2 7B architecture. It significantly improves language representation, particularly for low-resource languages, through continual pre-training on the extensive MaLA Corpus. This corpus comprises over 74 billion tokens across more than 500 languages.
Key Capabilities & Performance
- Massively Multilingual: Supports 546 languages, each with substantial training data (over 100k tokens).
- Diverse Task Proficiency: Excels in a wide range of tasks including commonsense reasoning, machine translation, text classification, natural language inference, code generation, and open-ended generation.
- Outperforms Llama 2-based models: Demonstrates superior performance in diverse multilingual settings, particularly in text classification and natural language inference.
- Intrinsic Evaluation: Achieves the lowest negative log-likelihood in intrinsic evaluations.
- Enhanced Code Generation: Shows improved performance in code generation and machine reading comprehension (MRC).
Training & Data
EMMA-500 leverages a diverse data mix including code, books, and instruction data. The model's development involved researchers from Helsinki-NLP, TU Darmstadt, the University of Edinburgh, and LMU Munich, funded by HPLT and UTTER.
Considerations
Challenges remain in low-resource languages, where the model may exhibit higher Self-BLEU scores, indicating reduced output diversity.