tartuNLP/Apertus-EstLLM-8B-Instruct-1125
The tartuNLP/Apertus-EstLLM-8B-Instruct-1125 is an 8 billion parameter instruction-following causal language model developed by TartuNLP and TalTechNLP research groups. It is fine-tuned from tartuNLP/Llama-3.1-EstLLM-8B-0525 and primarily focuses on Estonian and English languages. This model is intended as a research artifact, demonstrating minor improvements in Estonian benchmarks compared to its predecessor, but showing degradation in English performance.
Loading preview...
Model Overview
The tartuNLP/Apertus-EstLLM-8B-Instruct-1125 is an 8 billion parameter instruction-following causal language model, developed by the TartuNLP and TalTechNLP research groups with funding from the Estonian Ministry of Education and Research. It is based on the tartuNLP/Llama-3.1-EstLLM-8B-0525 model and is primarily designed for research purposes, not production use-cases.
Key Characteristics
- Multilingual Focus: Supports both Estonian and English, with a primary emphasis on enhancing Estonian language capabilities.
- Training Data: Continued pre-training involved a mix of Estonian National Corpus, Python-Edu, FineMath4-Plus, General Instruction-Augmented Corpora, and Cosmopedia v2. Supervised fine-tuning utilized approximately 764k examples, mainly from the Tulu 3 SFT mixture and EuroBlocks, with about 80% of examples in English. Direct Preference Optimization used English-only HelpSteer3.
- Performance: While showing minor improvements on Estonian benchmarks like IFEval-et and some Estonian language competence tasks compared to its direct predecessor, the model exhibits degradation in English instruction-following and knowledge/reasoning benchmarks when compared to other models of similar size.
- Context Length: Features a context length of 32768 tokens.
Intended Use
This model is explicitly designated as a research artifact and is not recommended for production environments due to its relatively small improvements in Estonian and observed degradation in English performance. It serves as a basis for further research into enhancing multilingual LLMs, particularly for less-resourced languages like Estonian.