SOURAV11318/Gemma-4-26B-Backup
Gemma 4-26B-Backup is a 26 billion parameter multimodal Mixture-of-Experts (MoE) model developed by Google DeepMind, part of the Gemma 4 family. It processes text, image, and video inputs, generating text outputs, and features a 256K token context window. This model is optimized for reasoning, coding, and agentic workflows, offering a balance of performance and efficient inference due to its active 3.8B parameters.
Loading preview...
Model Overview
SOURAV11318/Gemma-4-26B-Backup is a 26 billion parameter Mixture-of-Experts (MoE) model from the Gemma 4 family, developed by Google DeepMind. It is designed for multimodal understanding, capable of processing text, image, and video inputs to generate text outputs. This model features a substantial 256K token context window and supports over 140 languages.
Key Capabilities
- Multimodal Processing: Handles text, image, and video inputs, with interleaved multimodal input support.
- Reasoning: Incorporates configurable thinking modes for step-by-step problem-solving.
- Efficient Architecture: As an MoE model, it has 25.2B total parameters but only 3.8B active parameters, enabling faster inference comparable to a 4B model.
- Enhanced Coding & Agentic Capabilities: Shows significant improvements in coding benchmarks and includes native function-calling support for autonomous agents.
- Long Context: Supports a 256K token context window, utilizing a hybrid attention mechanism for efficient long-context processing.
Benchmark Highlights
The Gemma 4 26B A4B model demonstrates strong performance across various benchmarks, including:
- MMLU Pro: 82.6%
- AIME 2026 (no tools): 88.3%
- LiveCodeBench v6: 77.1%
- MMMU Pro (Vision): 73.8%
Intended Usage
This model is well-suited for a wide range of applications, including content creation (text generation, summarization), conversational AI, research, and educational tools. Its multimodal capabilities make it particularly effective for tasks involving image and video understanding, such as object detection, document parsing, and video analysis. The model's design also supports efficient deployment on consumer GPUs and workstations.