tartuNLP/Llama-3.1-EstLLM-8B-Instruct-1125

Hugging Face
TEXT GENERATIONPricing:Input $0.14 / Cached $0.028 / Output $0.26Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Nov 28, 2025License:llama3.1Architecture:Transformer0.0K Featherless Exclusive Warm

Llama-3.1-EstLLM-8B-Instruct-1125 is an instruction-following causal language model developed by TartuNLP and TalTechNLP, based on Meta's Llama 3.1 8B. It has undergone continuous pre-training on approximately 35 billion tokens, with a strong focus on Estonian language competence, knowledge, and reasoning. This model excels in Estonian language tasks, demonstrating significant improvements over its predecessor and competitive performance against larger models in specific Estonian benchmarks, while also maintaining strong English instruction-following capabilities.

Loading preview...

Model Overview

Llama-3.1-EstLLM-8B-Instruct-1125 is an instruction-following causal language model developed by the TartuNLP and TalTechNLP research groups. It is built upon Meta's Llama 3.1 8B, having undergone extensive continued pre-training on approximately 35 billion tokens, including significant Estonian language corpora, Python code, and mathematical datasets. This model further benefits from supervised fine-tuning (SFT) with around 764k examples and Direct Preference Optimization (DPO) using English-only HelpSteer3 data.

Key Capabilities and Differentiators

  • Enhanced Estonian Language Proficiency: Demonstrates substantial improvements in Estonian instruction-following, grammar, inflection, word meanings, and knowledge/reasoning benchmarks compared to its previous version and other models in its size class.
  • Bilingual Performance: While primarily focused on Estonian, it also shows strong performance in English instruction-following and competitive results in English knowledge and reasoning tasks like GSM8K.
  • Instruction-Following: Achieves a score of 0.6141 on IFEval-et (Estonian) and 0.8173 on IFEval-en (English), outperforming Llama-3.1-8B-Instruct in both.
  • Translation: Shows competitive BLEU scores for English to Estonian translation (0.2635 on wmt24pp).

Use Cases and Limitations

This model is particularly well-suited for applications requiring strong Estonian language understanding and generation, including chatbots, content creation, and language analysis in Estonian. It also performs well in English instruction-following scenarios. Current limitations include a relatively short context window of 4096 tokens, potential issues with multi-turn conversations, and a hard-coded date cut-off from the base Llama 3.1 system prompt.

Popular Sampler Settings

Top 3 parameter combinations used by Featherless users for this model. Click a tab to see each config.

temperature
top_p
top_k
frequency_penalty
presence_penalty
repetition_penalty
min_p