tartuNLP/Llama-3.1-EstLLM-70B-0826

Hugging Face
TEXT GENERATIONPricing:Input $3.5 / Cached $0.7 / Output $8.3Concurrent Unit Cost:4Model Size:70BQuant:FP8Context Size:8kTool Calling:SupportedPublished:Aug 17, 2026License:llama3.1Architecture:Transformer Featherless Exclusive Warm

The tartuNLP/Llama-3.1-EstLLM-70B-0826 is a 70 billion parameter causal language model developed by TartuNLP and TalTechNLP, continuously pre-trained on approximately 60 billion tokens from the Llama-3.1-70B base model. This model is specifically enhanced for Estonian language capabilities, demonstrating strong performance in Estonian grammar, inflection, and translation tasks. It is designed as a base model for further fine-tuning on downstream tasks, rather than direct use for chat or instruction-following.

Loading preview...

Model Overview

The tartuNLP/Llama-3.1-EstLLM-70B-0826 is a 70 billion parameter base causal language model, developed by the TartuNLP and TalTechNLP research groups with funding from the Estonian Ministry of Education and Research. It is built upon the meta-llama/Llama-3.1-70B architecture, having undergone continuous pre-training on an additional 60 billion tokens. This process has significantly enhanced its proficiency in both Estonian and English.

Key Capabilities and Performance

This model excels particularly in Estonian language tasks, outperforming its base Llama-3.1-70B model and other large models like Qwen2.5-72B in several key areas:

  • Estonian Language Proficiency: Demonstrates superior performance in grammar-et (0.8910) and inflection-et (0.9258), and achieves the highest score in trivia-et (0.5000) among the compared models.
  • Translation: Achieves strong BLEU scores for translation, with 29.73 for English-to-Estonian and 41.06 for Estonian-to-English, making it highly competitive for cross-lingual applications involving Estonian.
  • Base Model: It is provided as a base text completion model, intended for further fine-tuning to specific downstream tasks rather than direct instruction-following or chat applications.

When to Use This Model

  • Estonian-centric Applications: Ideal for developers and researchers focusing on applications requiring high proficiency in the Estonian language, such as content generation, analysis, or translation.
  • Fine-tuning Projects: Best suited as a foundation for fine-tuning on specialized datasets or tasks, leveraging its strong base capabilities in both Estonian and English.
  • Research and Development: Valuable for academic and industrial research into multilingual LLMs, particularly those exploring continued pre-training strategies for low-resource languages.