Goader/gemma-3-4b-it-uk-matt
Goader/gemma-3-4b-it-uk-matt is a 4.3 billion parameter instruction-tuned Gemma model developed by Goader, specifically adapted for Ukrainian language processing. It utilizes a Lapa tokenizer with MATT (Model-Aware Tokenizer Transfer) to significantly reduce token count for Ukrainian text, while retaining the original Gemma vocabulary and adding Ukrainian tokens. This model is optimized for efficient Ukrainian language generation and understanding, making it suitable for applications requiring strong performance in Ukrainian with reduced sequence length.
Loading preview...
Overview
Goader/gemma-3-4b-it-uk-matt is an instruction-tuned Gemma 3B model, developed by Goader, that has been transferred to a Ukrainian-centric Lapa tokenizer using MATT (Model-Aware Tokenizer Transfer). This process integrates Ukrainian tokens into the original Gemma vocabulary, reducing the token count for Ukrainian text by approximately one-third. The model's input embeddings were initialized using FOCUS and then trained against the frozen original model over 1.03 million Ukrainian documents from the Kobza corpus.
Key Capabilities
- Efficient Ukrainian Processing: Achieves roughly a 35% reduction in token count for Ukrainian sequences compared to the original Gemma model.
- Strong Ukrainian Generation: Recovers 97-99% of the original model's performance on Ukrainian generation tasks (e.g., FLORES en→uk).
- Ukrainian Comprehension: Recovers 82-91% of the original model's accuracy on understanding-heavy Ukrainian tasks, balancing comprehension with sequence length efficiency.
Good For
- Applications requiring efficient and accurate text generation in Ukrainian.
- Use cases where reducing token count for Ukrainian input/output is beneficial.
- Instruction-following tasks in Ukrainian, leveraging the base Gemma-3-4b-it capabilities with enhanced Ukrainian support.