google/gemma-4-26B-A4B-it

Hugging Face
TEXT GENERATIONConcurrent Unit Cost:2Model Size:26BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Mar 11, 2026License:apache-2.0Architecture:Transformer1.4K Open Weights Warm

The google/gemma-4-26B-A4B-it model is a multimodal, instruction-tuned language model from Google DeepMind, part of the Gemma 4 family. This 25.2 billion parameter Mixture-of-Experts (MoE) model, with 3.8 billion active parameters, processes text and image inputs, generating text outputs. It features a 256K token context window and is optimized for reasoning, agentic workflows, and coding, offering fast inference due to its efficient MoE architecture.

Loading preview...

Gemma 4 26B A4B MoE: Multimodal Instruction-Tuned Model

This model is a member of the Gemma 4 family, developed by Google DeepMind, offering advanced multimodal capabilities. It is an instruction-tuned variant designed to handle both text and image inputs, producing text outputs. A key feature is its Mixture-of-Experts (MoE) architecture, which, despite having 25.2 billion total parameters, utilizes only 3.8 billion active parameters during inference. This design allows for significantly faster inference speeds, comparable to a 4B parameter model, making it efficient for deployment.

Key Capabilities

  • Multimodal Processing: Supports text and image inputs, with variable aspect ratio and resolution for images. Video processing is also supported by analyzing sequences of frames.
  • Extended Context Window: Features a substantial 256K token context window, enabling deep awareness for complex, long-context tasks.
  • Reasoning: Designed with configurable thinking modes to enhance reasoning capabilities, allowing the model to process information step-by-step.
  • Enhanced Coding & Agentic Capabilities: Demonstrates improvements in coding benchmarks and includes native function-calling support for autonomous agents.
  • Native System Prompt Support: Integrates native support for the system role, facilitating more structured and controllable conversations.

Good For

  • Fast Multimodal Inference: Its MoE architecture makes it suitable for applications requiring quick responses from multimodal inputs.
  • Complex Reasoning Tasks: The model's reasoning capabilities and large context window are beneficial for intricate problem-solving.
  • Agentic Workflows and Coding: Excels in scenarios requiring code generation, completion, correction, and structured tool use.
  • Image Understanding: Capable of object detection, document parsing, chart comprehension, and OCR, supporting diverse visual analysis tasks.

Popular Sampler Settings

Top 3 parameter combinations used by Featherless users for this model. Click a tab to see each config.

temperature
top_p
top_k
frequency_penalty
presence_penalty
repetition_penalty
min_p