ananddey/akhorika-e2b-base

VISIONConcurrent Unit Cost:1Model Size:5.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 22, 2026License:gemmaArchitecture:Transformer Featherless Exclusive Cold

Akhorika E2B Base by ananddey is a 2.69 billion parameter (1.91B active text decoder) language model, continued-pretrained on Assamese language corpora. Based on Google's Gemma 4 E2B architecture, it features a custom Assamese-first SentencePiece tokenizer and FOCUS embedding initialization. This model is specifically designed for generative tasks in Assamese and English, demonstrating a perplexity of approximately 60 on its held-out Assamese evaluation set.

Loading preview...

Akhorika E2B Base Overview

Akhorika E2B Base is a 2.69 billion parameter language model developed by ananddey, with 1.91 billion active text decoder parameters. It is built upon Google's Gemma 4 E2B architecture and has undergone continued pretraining (CPT) using Assamese language corpora. This model is distinguished by its specialized focus on the Assamese language, alongside English.

Key Capabilities & Features

  • Architecture: Based on Gemma4ForConditionalGeneration (Google Gemma 4 E2B).
  • Parameter Count: 2.69B total parameters, with 1.91B dedicated to the text decoder.
  • Tokenizer: Utilizes a custom 32,000 vocabulary Assamese-first SentencePiece Unigram tokenizer, replacing the native 256k tokenizer for optimized performance in Assamese.
  • Embedding Initialization: Employs FOCUS (semantic subword surface-form projection from base Gemma embeddings) for enhanced linguistic representation.
  • Language Support: Primarily supports Assamese (as) and English (en).
  • Training: Continued Pretraining (CPT) completed over one epoch, reaching step 821.
  • Performance: Achieved an evaluation loss of 4.09, corresponding to a perplexity of approximately 60 on the held-out Assamese evaluation dataset.

Good For

  • Assamese Language Generation: Ideal for applications requiring text generation, translation, or understanding in Assamese.
  • Bilingual Applications: Suitable for use cases involving both Assamese and English content.
  • Research & Development: Provides a strong foundation for further fine-tuning or research into low-resource language models, particularly for Assamese.