tartuNLP/Llama-3.1-EstLLM-8B-Instruct-1125
The tartuNLP/Llama-3.1-EstLLM-8B-Instruct-1125 is an 8 billion parameter instruction-following causal language model developed by TartuNLP and TalTechNLP. It is based on Meta's Llama 3.1 architecture and has been continuously pre-trained on approximately 35 billion tokens, with further supervised fine-tuning and direct preference optimization. This model excels in Estonian language competence, instruction-following, and knowledge-based tasks, demonstrating significant improvements over its previous version and strong performance in English benchmarks.
Loading preview...
Model Overview
tartuNLP/Llama-3.1-EstLLM-8B-Instruct-1125 is an 8 billion parameter instruction-following model developed by the TartuNLP and TalTechNLP research groups, funded by the Estonian Ministry of Education and Research. It is built upon Meta's Llama 3.1-8B and underwent extensive continued pre-training on approximately 35 billion tokens, including a significant portion of Estonian National Corpus, Python-Edu, FineMath4-Plus, and general instruction-augmented corpora. This was followed by supervised fine-tuning using 764k examples from datasets like Tulu 3 SFT mixture and EuroBlocks-SFT-Synthetic, with additional data from the Institute of Estonian Language (EKI).
Key Capabilities & Performance
This model demonstrates strong performance in both Estonian and English language tasks, particularly in instruction-following and multiple-choice evaluations. It shows notable improvements over its predecessor, Llama-3.1-EstLLM-8B-Instruct-0825, across various benchmarks. For instance, it achieved 0.6141 on IFEval-et (Estonian instruction-following) and 0.8173 on IFEval-en (English instruction-following), surpassing Llama-3.1-8B-Instruct in English. In Estonian language competence, it scored 0.831 on Grammar-et and 0.9619 on Word-Meanings-et. It also shows competitive results in English knowledge and reasoning benchmarks like GSM8K and MMLU-Redux.
Good For
- Estonian Language Applications: Excels in Estonian instruction-following, grammar, inflection, and word meaning tasks, making it highly suitable for applications requiring strong Estonian language understanding and generation.
- Bilingual (Estonian/English) Use Cases: Offers robust performance in both Estonian and English, making it valuable for bilingual applications or tasks involving translation from English to Estonian.
- Instruction Following: Designed for instruction-following tasks, providing reliable responses based on given prompts.
Limitations
As an early prototype, it has a relatively short context of 4096 tokens, which may limit performance on longer contexts. While improved by merging, multi-turn conversations are not fully guaranteed, and it inherits the base Llama 3.1 system prompt's hard-coded date cut-off.
Top 3 parameter combinations used by Featherless users for this model. Click a tab to see each config.