RedHatAI/gemma-4-26B-A4B-it

VISIONConcurrent Unit Cost:2Model Size:26BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 14, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

RedHatAI/gemma-4-26B-A4B-it is an instruction-tuned multimodal language model from the Gemma 4 family, developed by Google DeepMind. This Mixture-of-Experts (MoE) model features 25.2 billion total parameters with 3.8 billion active parameters, a 256K token context window, and supports text and image inputs. It is optimized for fast inference, excelling in reasoning, agentic workflows, coding, and multimodal understanding tasks.

Loading preview...

Gemma 4 26B A4B-it: Multimodal MoE for Advanced AI

RedHatAI/gemma-4-26B-A4B-it is an instruction-tuned model from Google DeepMind's Gemma 4 family, designed for frontier-level performance across various AI tasks. This particular variant utilizes a Mixture-of-Experts (MoE) architecture, featuring 25.2 billion total parameters but only 3.8 billion active parameters during inference, allowing for significantly faster execution compared to dense models of similar scale. It supports a substantial 256K token context window and is multimodal, processing both text and image inputs.

Key Capabilities

  • Multimodality: Handles interleaved text and image inputs, with support for variable aspect ratios and resolutions, object detection, document parsing, and OCR.
  • Reasoning: Incorporates a built-in reasoning mode for step-by-step thought processes before generating answers.
  • Coding & Agentic Workflows: Shows enhanced performance in coding benchmarks and includes native function-calling support for autonomous agents.
  • Long Context: Features a 256K token context window, enabling deep awareness for complex, long-context tasks through a hybrid attention mechanism.
  • Multilingual Support: Pre-trained on over 140 languages, with out-of-the-box support for 35+ languages.

Good For

  • Fast Inference: Its MoE architecture makes it suitable for applications requiring high-speed processing without sacrificing capability.
  • Complex Multimodal Tasks: Ideal for scenarios involving detailed image understanding combined with textual analysis, such as document processing or visual question answering.
  • Advanced Reasoning & Coding: Benefits applications needing strong logical reasoning, code generation, or agentic behaviors.
  • Structured Conversations: Native system prompt support allows for more controllable and structured conversational AI.