alphaedge-ai/gemma-3-4b-it-eng-32768
VISIONPricing:Input $0.2 / Output $0.4Concurrent Unit Cost:1Model Size:4.3BQuant:BF16Context Size:32kPublished:May 6, 2026License:gemmaArchitecture:Transformer0.0K Featherless Exclusive Cold
alphaedge-ai/gemma-3-4b-it-eng-32768 is a 4.3 billion parameter instruction-tuned Gemma model, derived from google/gemma-3-4b-it. It is specifically optimized for English language tasks through an 87.50% vocabulary size reduction to 32,768 tokens, resulting in a 13.66% smaller model size. This trimming technique aims to maintain performance for English while significantly reducing memory footprint, making it suitable for efficient English-centric applications.
Loading preview...
Model Overview
This model, alphaedge-ai/gemma-3-4b-it-eng-32768, is a specialized version of the google/gemma-3-4b-it instruction-tuned model. It has been optimized for English language performance by significantly reducing its vocabulary size using a technique called trimming.
Key Characteristics
- Reduced Size: The model is 13.66% smaller than its original counterpart, with its parameter count reduced from 4.3 billion to approximately 3.7 billion.
- Optimized Vocabulary: The vocabulary size has been drastically cut by 87.50%, from 262,144 tokens to 32,768 tokens. This reduction focuses on retaining tokens most relevant to the English language.
- Memory Efficiency: The trimming process leads to a much smaller memory footprint, making it more efficient for deployment in English-only contexts.
- English-Centric: While designed to perform similarly to the original model for English, its performance for other languages may be degraded due to the removal of non-English-centric tokens.
Ideal Use Cases
- Applications requiring a smaller, more efficient Gemma-3B model for English language processing.
- Scenarios where memory footprint is a critical constraint and the primary language is English.
- Developers looking for a Gemma-based model with optimized inference for English-only tasks.