google/gemma-4-26B-A4B

Hugging Face
TEXT GENERATIONConcurrent Unit Cost:2Model Size:26BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Mar 12, 2026License:apache-2.0Architecture:Transformer0.3K Open Weights Warm

google/gemma-4-26B-A4B is a 25.2 billion total parameter Mixture-of-Experts (MoE) multimodal model developed by Google DeepMind, part of the Gemma 4 family. It processes text and image inputs, generating text outputs, and features 3.8 billion active parameters for efficient inference. With a 256K token context window, it excels in reasoning, coding, and agentic workflows, making it suitable for consumer GPUs and workstations.

Loading preview...

Overview of Gemma-4-26B-A4B

The google/gemma-4-26B-A4B is a multimodal model from Google DeepMind's Gemma 4 family, designed to handle both text and image inputs and generate text outputs. This specific variant is a Mixture-of-Experts (MoE) model with 25.2 billion total parameters, but it operates with only 3.8 billion active parameters during inference, allowing for faster execution compared to dense models of similar total size. It supports a substantial 256K token context window and is optimized for reasoning, coding, and agentic capabilities.

Key Capabilities

  • Multimodal Input: Processes text and images, with support for variable aspect ratios and resolutions.
  • Reasoning: Designed with configurable thinking modes for enhanced reasoning.
  • Efficient Architecture: Utilizes a Mixture-of-Experts (MoE) design for optimized inference speed.
  • Extended Context: Features a 256K token context window for complex, long-context tasks.
  • Enhanced Coding & Agentic Capabilities: Achieves notable improvements in coding benchmarks and includes native function-calling support.
  • Native System Prompt Support: Integrates native support for the system role for more structured conversations.

Good For

  • Complex Reasoning Tasks: Its design and thinking modes make it suitable for tasks requiring deep logical processing.
  • Coding and Agentic Workflows: Strong performance in coding benchmarks and function-calling support make it ideal for code generation and autonomous agents.
  • Multimodal Understanding: Capable of image understanding, including object detection, document parsing, and chart comprehension.
  • Deployment on Consumer Hardware: The efficient MoE architecture allows for deployment on consumer GPUs and workstations, balancing performance with accessibility.

Popular Sampler Settings

Top 3 parameter combinations used by Featherless users for this model. Click a tab to see each config.

temperature
top_p
top_k
frequency_penalty
presence_penalty
repetition_penalty
min_p