COPA-AI/arm-gemma-e4b
COPA-AI/arm-gemma-e4b is a 7.9 billion parameter base model, adapted from Gemma-4-E4B, specifically for the Armenian language through continued pretraining. Developed by COPA-AI, it is the first open Armenian LLM released with its complete training corpus and recipe, excelling in Armenian text continuation and likelihood scoring. This model demonstrates superior performance across a six-task Armenian likelihood suite and various generative tasks, making it ideal for Armenian NLP applications.
Loading preview...
arm-gemma-e4b: The First Open Armenian LLM
arm-gemma-e4b is a 7.9 billion parameter base model, adapted from Gemma-4-E4B, and represents the first open Armenian Large Language Model released with its full training corpus and recipe. Developed by COPA-AI, this model underwent continued pretraining on 10 billion tokens, primarily from Armenian web and STEM data, alongside English web and code.
Key Capabilities & Features
- Armenian Language Specialization: Achieves the highest six-task mean among evaluated open Armenian models (0.499), outperforming prior models and even the unadapted base Gemma-4-E4B.
- Comprehensive Training Data: Trained on a unique mixture including 69% ArmWeb, 4% ArmSTEM-HY, 2% ArmSTEM-EN, 20% English web replay, and 5% code, with every training token either public or reproducible.
- Robust Performance: Shows significant improvements in Armenian likelihood tasks like Belebele-hye (+0.097) and INCLUDE-Armenian (+0.040), and dramatic gains in generative tasks such as SynDARin (0.04 to 0.92) and Hartak (0.02 to 0.82).
- Base Model: Designed for Armenian text continuation, likelihood scoring, or as a foundation for further instruction tuning, without instruction tuning or a chat template.
- Data Decontamination: Training data was rigorously decontaminated against ten Armenian evaluation sets and ArmBench items to ensure benchmark integrity.
Use Cases
- Armenian Text Generation: Ideal for generating coherent and contextually relevant Armenian text.
- Likelihood Scoring: Useful for evaluating the probability of Armenian text sequences.
- Foundation for Fine-tuning: Serves as an excellent starting point for developing instruction-tuned or chat-enabled Armenian LLMs.
- Research & Development: Provides a transparent and reproducible platform for research into low-resource language LLMs, particularly for Armenian.