INSAIT-Institute/BgGPT-Gemma-2-27B-IT-v1.0

Hugging Face
TEXT GENERATIONConcurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kPublished:Nov 15, 2024License:gemmaArchitecture:Transformer0.0K Featherless Exclusive Warm

INSAIT-Institute/BgGPT-Gemma-2-27B-IT-v1.0 is a 27 billion parameter Bulgarian language model developed by INSAIT, based on Google's Gemma 2 architecture with a 32768 token context length. It was continuously pre-trained on 100 billion tokens, primarily Bulgarian, using a Branch-and-Merge strategy to achieve outstanding Bulgarian cultural and linguistic capabilities while retaining English performance. This model excels in Bulgarian language understanding, logical reasoning, and chat performance, outperforming much larger models in Bulgarian benchmarks.

Loading preview...

BgGPT-Gemma-2-27B-IT-v1.0: A Specialized Bulgarian LLM

Developed by INSAIT, BgGPT-Gemma-2-27B-IT-v1.0 is a 27 billion parameter instruction-tuned language model built upon Google's Gemma 2 architecture. This model is specifically designed to excel in the Bulgarian language, having undergone continuous pre-training on approximately 100 billion tokens, with 85 billion being Bulgarian. This extensive training, utilizing a Branch-and-Merge strategy, ensures strong Bulgarian linguistic and cultural understanding without compromising its inherited English performance.

Key Capabilities & Performance

  • Superior Bulgarian Performance: Achieves excellent results across various Bulgarian benchmarks, including Winogrande, Hellaswag, ARC, TriviaQA, GSM-8k, Exams, and MON, often outperforming significantly larger models like Alibaba's Qwen 2.5 72B and Meta's Llama3.1 70B.
  • Enhanced Chat Performance: Demonstrates chat performance in Bulgarian that surpasses smaller commercial models (e.g., Claude Haiku, GPT-4o-mini) and is on par with leading commercial models (e.g., Claude Sonnet, GPT-4o), as evaluated on thousands of real-world Bulgarian conversations.
  • Retained English Proficiency: Maintains the strong English language capabilities of the original Google Gemma 2 models.
  • Instruction-Tuned: Fine-tuned on a newly constructed Bulgarian instruction dataset derived from real-world conversations.

Ideal Use Cases

  • Bulgarian Language Applications: Excellent for tasks requiring deep understanding and generation in Bulgarian, including content creation, customer support, and educational tools.
  • Multilingual Environments: Suitable for applications needing strong performance in both Bulgarian and English.
  • Research and Development: Provides a powerful base for further research into low-resource language modeling and cross-lingual transfer.