tartuNLP/Apertus-EstLLM-8B-Instruct-1125

TEXT GENERATIONPricing:Input $0.431 / Cached $0.0216 / Output $1.12Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kPublished:Jan 23, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

The tartuNLP/Apertus-EstLLM-8B-Instruct-1125 is an 8 billion parameter instruction-following causal language model developed by TartuNLP and TalTechNLP research groups. It is fine-tuned from tartuNLP/Llama-3.1-EstLLM-8B-0525 and primarily focuses on Estonian and English languages. This model is intended as a research artifact, demonstrating minor improvements in Estonian benchmarks compared to its predecessor, but showing degradation in English performance.

Loading preview...

Model Overview

The tartuNLP/Apertus-EstLLM-8B-Instruct-1125 is an 8 billion parameter instruction-following causal language model, developed by the TartuNLP and TalTechNLP research groups with funding from the Estonian Ministry of Education and Research. It is based on the tartuNLP/Llama-3.1-EstLLM-8B-0525 model and is primarily designed for research purposes, not production use-cases.

Key Characteristics

  • Multilingual Focus: Supports both Estonian and English, with a primary emphasis on enhancing Estonian language capabilities.
  • Training Data: Continued pre-training involved a mix of Estonian National Corpus, Python-Edu, FineMath4-Plus, General Instruction-Augmented Corpora, and Cosmopedia v2. Supervised fine-tuning utilized approximately 764k examples, mainly from the Tulu 3 SFT mixture and EuroBlocks, with about 80% of examples in English. Direct Preference Optimization used English-only HelpSteer3.
  • Performance: While showing minor improvements on Estonian benchmarks like IFEval-et and some Estonian language competence tasks compared to its direct predecessor, the model exhibits degradation in English instruction-following and knowledge/reasoning benchmarks when compared to other models of similar size.
  • Context Length: Features a context length of 32768 tokens.

Intended Use

This model is explicitly designated as a research artifact and is not recommended for production environments due to its relatively small improvements in Estonian and observed degradation in English performance. It serves as a basis for further research into enhancing multilingual LLMs, particularly for less-resourced languages like Estonian.