Vegss/gemma-4-26B-A4B

VISIONConcurrent Unit Cost:2Model Size:26BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 8, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The Gemma 4 26B A4B model by Google DeepMind is a 25.2 billion parameter multimodal Mixture-of-Experts (MoE) model, featuring 3.8 billion active parameters for efficient inference. It processes text, image, and video inputs, generating text outputs, and supports a 256K token context window. This model excels in reasoning, agentic workflows, and coding, offering a balance of performance and speed for consumer GPUs and workstations.

Loading preview...

Gemma 4 26B A4B: Multimodal MoE for Advanced Reasoning and Coding

Developed by Google DeepMind, Gemma 4 is a family of open multimodal models, with the 26B A4B variant being a Mixture-of-Experts (MoE) architecture. This model handles text, image, and video inputs, producing text outputs, and supports a substantial 256K token context window. Its MoE design, with 25.2 billion total parameters and 3.8 billion active parameters, allows it to run efficiently, almost as fast as a 4B model, making it suitable for consumer GPUs and workstations.

Key Capabilities

  • Multimodality: Processes text, image, and video inputs, with variable aspect ratio and resolution support for images.
  • Reasoning: Designed with configurable thinking modes for highly capable reasoning.
  • Agentic & Coding: Enhanced coding benchmarks and native function-calling support for autonomous agents.
  • Long Context: Features a 256K token context window, utilizing a hybrid attention mechanism for efficiency.
  • Native System Prompt Support: Enables more structured and controllable conversations.

Good For

  • Reasoning-intensive tasks: Leverages built-in reasoning modes for complex problem-solving.
  • Agentic workflows: Supports function calling for building autonomous agents.
  • Coding: Excels in code generation, completion, and correction.
  • Multimodal understanding: Ideal for applications requiring analysis of combined text, image, and video data.
  • Efficient deployment: Its MoE architecture provides high performance with efficient inference, suitable for various hardware.