RedHatAI/gemma-4-26B-A4B

VISIONConcurrent Unit Cost:2Model Size:26BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 15, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The RedHatAI/gemma-4-26B-A4B is a 26 billion parameter multimodal Mixture-of-Experts (MoE) model developed by Google DeepMind, part of the Gemma 4 family. It processes text and image inputs, generating text outputs, and features an active parameter count of 3.8 billion for efficient inference. This model excels in reasoning, coding, and agentic workflows, supporting a 256K token context window and multilingual capabilities across 140+ languages.

Loading preview...

Gemma 4 26B A4B: Multimodal MoE for Reasoning and Coding

This model is a 26 billion parameter Mixture-of-Experts (MoE) variant from the Gemma 4 family, developed by Google DeepMind. It is designed for efficient inference with only 3.8 billion active parameters, making it perform comparably to a 4B model in speed while leveraging its larger total parameter count for capability. The model supports a substantial 256K token context window and is multimodal, handling both text and image inputs to generate text outputs.

Key Capabilities

  • Multimodality: Processes text and images, with variable aspect ratio and resolution support. Video understanding is also supported by processing sequences of frames.
  • Efficient Architecture: Utilizes a Mixture-of-Experts (MoE) design, activating a smaller subset of parameters (3.8B) during inference for speed.
  • Extended Context Window: Features a 256K token context length, enabling deep awareness for complex, long-context tasks through a hybrid attention mechanism.
  • Enhanced Reasoning & Coding: Designed as a highly capable reasoner with configurable thinking modes and significant improvements in coding benchmarks, including native function-calling support.
  • Multilingual Support: Pre-trained on over 140 languages, with out-of-the-box support for 35+ languages.
  • Native System Prompt Support: Introduces native support for the system role for more structured and controllable conversations.

Benchmark Highlights

The 26B A4B MoE model demonstrates strong performance across various benchmarks:

  • MMLU Pro: 82.6%
  • AIME 2026 no tools: 88.3%
  • LiveCodeBench v6: 77.1%
  • Codeforces ELO: 1718
  • MMMU Pro (Vision): 73.8%

Good For

  • Agentic Workflows: Enhanced coding and function-calling capabilities make it suitable for building autonomous agents.
  • Complex Reasoning Tasks: Designed for high-capability reasoning with configurable thinking modes.
  • Code Generation and Understanding: Achieves notable improvements in coding benchmarks.
  • Multimodal Applications: Ideal for applications requiring both text and image understanding, such as document parsing, UI understanding, and OCR.
  • Long Context Processing: Its 256K token context window is beneficial for tasks requiring extensive contextual awareness.