google/gemma-4-26B-A4B-it-qat-q4_0-unquantized

Hugging Face
VISIONConcurrent Unit Cost:2Model Size:26BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Apr 29, 2026License:apache-2.0Architecture:Transformer0.1K Open Weights Featherless Exclusive Warm

The google/gemma-4-26B-A4B-it-qat-q4_0-unquantized model is a multimodal instruction-tuned variant from the Gemma 4 family, developed by Google DeepMind. This 25.2 billion total parameter Mixture-of-Experts (MoE) model, with 3.8 billion active parameters, is optimized with Quantization-Aware Training (QAT) for efficient deployment while maintaining quality. It features a 256K token context window and excels in reasoning, coding, and multimodal understanding across text and image inputs.

Loading preview...

Gemma 4 26B A4B MoE: QAT-Optimized Multimodal LLM

This model is part of the Gemma 4 family by Google DeepMind, specifically an instruction-tuned variant optimized with Quantization-Aware Training (QAT) for reduced memory footprint while preserving performance. It's a Mixture-of-Experts (MoE) architecture with 25.2 billion total parameters, but only 3.8 billion active parameters, enabling faster inference compared to dense models of similar total size. The model supports a substantial 256K token context window and is designed for multimodal tasks, processing both text and image inputs.

Key Capabilities

  • Multimodal Understanding: Processes text and image inputs, with support for variable aspect ratios and resolutions. It can analyze images for object detection, document parsing, OCR, and more.
  • Reasoning: Designed as a highly capable reasoner with configurable thinking modes.
  • Coding & Agentic Capabilities: Achieves notable improvements in coding benchmarks and includes native function-calling support for autonomous agents.
  • Long Context: Features a 256K token context window, suitable for complex, long-context tasks.
  • Efficient Architecture: The MoE design allows for fast inference by activating only a subset of parameters.
  • Native System Prompt Support: Introduces native support for the system role for more structured conversations.

Good for

  • Applications requiring efficient multimodal processing (text and image) with a large context window.
  • Tasks demanding strong reasoning and coding abilities, including agentic workflows.
  • Deployment scenarios where memory efficiency and faster inference are critical, leveraging its QAT optimization and MoE architecture.