GiorgiGE/Kolkha-Mini-Georgian
Kolkha-Mini is a 2 billion parameter causal language model developed by GiorgiGE, fine-tuned from Qwen/Qwen3-1.7B to specialize in the Georgian language. This model prioritizes coherent Georgian text generation and serves as an early-stage foundation for Georgian-focused NLP research and further fine-tuning. It offers a 32768 token context length and is optimized for low-resource language modeling.
Loading preview...
Kolkha-Mini: A Georgian Language Foundation Model
Kolkha-Mini is a 2 billion parameter language model developed by GiorgiGE, specifically fine-tuned from the Qwen/Qwen3-1.7B base model to specialize in the Georgian language. It was trained using QLoRA (4-bit) for causal language modeling over 2 epochs, with a context length of 1024 tokens during training, resulting in a fully merged FP16 model.
Key Capabilities
- Produces coherent Georgian text.
- Understands Georgian sentence structure.
- Serves as a solid starting point for further fine-tuning and research.
Intended Use Cases
- Georgian language research.
- Further fine-tuning for specific applications.
- Dataset experimentation.
- Low-resource language modeling.
While it excels at generating Georgian text, it is important to note its current limitations, including occasional grammatical inaccuracies, hallucinations, and invented words. It is not instruction-tuned or safety-aligned and is not recommended for production deployment or high-stakes factual tasks. Kolkha-Mini is designed as a base to build upon, with performance expected to improve significantly with larger and cleaner datasets.