lapa-llm/lapa-v0.1.2-instruct

Hugging Face
VISIONPricing:Input $0.2 / Output $0.6Concurrent Unit Cost:1Model Size:12BQuant:FP8Context Size:32kPublished:Oct 18, 2025License:gemmaArchitecture:Transformer0.0K Featherless Exclusive Warm

lapa-llm/lapa-v0.1.2-instruct is a 12 billion parameter instruction-tuned causal language model developed by a team of Ukrainian researchers, based on Gemma-3-12B. It features a highly optimized tokenizer for Ukrainian, requiring 1.5 times fewer tokens for Ukrainian text compared to the original Gemma 3, making it efficient for Ukrainian language processing. This model excels in Ukrainian-to-English translation, image processing in Ukrainian, summarization, and Q&A tasks, making it suitable for RAG systems and culturally-aware Ukrainian text generation.

Loading preview...

Lapa LLM v0.1.2-instruct: Optimized for Ukrainian Language Processing

Lapa LLM v0.1.2-instruct is a 12 billion parameter open-source language model, built upon Gemma-3-12B, developed by a collaborative team of Ukrainian researchers. Its primary focus is on highly efficient and accurate Ukrainian language processing. The model is named in honor of Valentyn Lapa, a pioneer in data handling methods.

Key Capabilities & Differentiators

  • Best-in-class Ukrainian Tokenizer: Features a state-of-the-art tokenizer adapted for Ukrainian, reducing token count by 1.5x compared to the original Gemma 3 for Ukrainian text, leading to faster and more efficient processing.
  • High Performance on Ukrainian Benchmarks: Achieves strong results in instruction-tuned benchmarks, closely trailing leading models like MamayLM in some categories. It is a leader in pretraining benchmarks for Ukrainian.
  • Multimodal Support: One of the best models in its size class for image processing in Ukrainian, as measured on the MMZNO benchmark.
  • Excellent Translation & RAG: Demonstrates the best English-to-Ukrainian and Ukrainian-to-English translation quality (33 BLEU on FLORES) and strong performance in summarization and Q&A, making it ideal for RAG systems.
  • Maximum Openness: The project emphasizes transparency, providing the model for commercial use, publishing approximately 25 training datasets, disclosing data filtering methods (including for disinformation), and offering open-source code and training documentation.

Ideal Use Cases

  • Sensitive Document Processing: Enables local processing of sensitive Ukrainian documents without external server transfers.
  • Culturally Aware Text Generation: Generates Ukrainian text that respects cultural and historical context, avoiding code-switching.
  • RAG Systems & Chatbots: Building robust RAG systems and chatbots that produce high-quality Ukrainian output.
  • Specialized Solutions: Fine-tuning for specific Ukrainian NLP tasks.
  • Machine Translation: Achieving top-tier translation quality between English and Ukrainian.