ananddey/akhorika-e2b-base
Akhorika E2B Base by ananddey is a 2.69 billion parameter (1.91B active text decoder) language model, continued-pretrained on Assamese language corpora. Based on Google's Gemma 4 E2B architecture, it features a custom Assamese-first SentencePiece tokenizer and FOCUS embedding initialization. This model is specifically designed for generative tasks in Assamese and English, demonstrating a perplexity of approximately 60 on its held-out Assamese evaluation set.
Loading preview...
Akhorika E2B Base Overview
Akhorika E2B Base is a 2.69 billion parameter language model developed by ananddey, with 1.91 billion active text decoder parameters. It is built upon Google's Gemma 4 E2B architecture and has undergone continued pretraining (CPT) using Assamese language corpora. This model is distinguished by its specialized focus on the Assamese language, alongside English.
Key Capabilities & Features
- Architecture: Based on
Gemma4ForConditionalGeneration(Google Gemma 4 E2B). - Parameter Count: 2.69B total parameters, with 1.91B dedicated to the text decoder.
- Tokenizer: Utilizes a custom 32,000 vocabulary Assamese-first SentencePiece Unigram tokenizer, replacing the native 256k tokenizer for optimized performance in Assamese.
- Embedding Initialization: Employs FOCUS (semantic subword surface-form projection from base Gemma embeddings) for enhanced linguistic representation.
- Language Support: Primarily supports Assamese (
as) and English (en). - Training: Continued Pretraining (CPT) completed over one epoch, reaching step 821.
- Performance: Achieved an evaluation loss of 4.09, corresponding to a perplexity of approximately 60 on the held-out Assamese evaluation dataset.
Good For
- Assamese Language Generation: Ideal for applications requiring text generation, translation, or understanding in Assamese.
- Bilingual Applications: Suitable for use cases involving both Assamese and English content.
- Research & Development: Provides a strong foundation for further fine-tuning or research into low-resource language models, particularly for Assamese.