majentik/gemma-4-E4B-turboquant

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Apr 6, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

The majentik/gemma-4-E4B-turboquant model is a documentation repository for applying TurboQuant KV cache compression to the google/gemma-4-E4B base model, a 4-billion parameter dense transformer with a 128K context length. This model focuses on reducing memory usage for the attention cache during inference, which is critical for long context processing. It supports text, image, and audio modalities and is designed for edge-optimized applications, offering runtime memory savings without modifying the base model's weights.

Loading preview...

Overview

majentik/gemma-4-E4B-turboquant is a documentation repository detailing how to integrate TurboQuant KV cache compression with the google/gemma-4-E4B base model. This approach significantly reduces the memory footprint of the attention cache during inference, which is particularly beneficial for processing long contexts. Unlike weight quantization, KV cache compression is applied at runtime, allowing the same base model weights to be used with or without this optimization.

Key Capabilities

  • KV Cache Compression: Reduces attention memory usage at inference time, crucial for extended context lengths.
  • Runtime Application: Applied dynamically, meaning it's compatible with existing google/gemma-4-E4B weights.
  • Hardware Agnostic: Can be paired with any weight variant of the base model, providing runtime savings across various devices.
  • Multi-modal Support: The underlying gemma-4-E4B model supports text, image, and audio modalities.

When to Use This Model

This model is ideal for developers looking to:

  • Optimize Memory Usage: Especially when working with the 128K context length of gemma-4-E4B.
  • Improve Inference Efficiency: By reducing the KV cache memory, it can enable longer sequences or larger batch sizes on constrained hardware.
  • Leverage Existing Weights: Apply compression without needing to re-quantize or modify the base model's weights.

While the original TurboQuant llama.cpp fork is now considered legacy, upstream llama.cpp and Ollama offer native KV cache quantization options (q8_0, q4_0) that provide similar memory benefits.