AlexHung29629/gemma-4-26B-A4B-it-agent
The AlexHung29629/gemma-4-26B-A4B-it-agent is a 25.2 billion parameter Mixture-of-Experts (MoE) multimodal model from the Gemma 4 family by Google DeepMind, featuring 3.8 billion active parameters and a 256K token context window. This instruction-tuned agent model processes text, image, and video inputs to generate text outputs, excelling in reasoning, coding, and agentic workflows. Its hybrid attention mechanism and MoE architecture optimize for fast inference and deep awareness in complex, long-context tasks.
Loading preview...
Overview
AlexHung29629/gemma-4-26B-A4B-it-agent is a 25.2 billion parameter instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind's Gemma 4 family. It features 3.8 billion active parameters, allowing for fast inference comparable to a 4B model, and supports a 256K token context window. This multimodal model handles text, image, and video inputs, generating text outputs, and is designed for high-performance on consumer GPUs and workstations.
Key Capabilities
- Multimodal Processing: Supports interleaved text, image, and video inputs with variable aspect ratio and resolution for images.
- Reasoning & Agentic Workflows: Designed as a highly capable reasoner with configurable thinking modes and native function-calling support for autonomous agents.
- Coding: Achieves notable improvements in coding benchmarks, supporting code generation, completion, and correction.
- Long Context: Features a 256K token context window, utilizing a hybrid attention mechanism for efficient long-context processing.
- Native System Prompt Support: Enhances structured and controllable conversations.
Good for
- Complex Reasoning Tasks: Its built-in reasoning mode and strong performance on benchmarks like MMLU Pro (82.6%) make it suitable for intricate problem-solving.
- Agentic Applications: Native function-calling and agentic capabilities are ideal for developing sophisticated AI agents.
- Coding Assistance: Excels in code generation and related tasks, as indicated by its LiveCodeBench v6 score of 77.1%.
- Multimodal Understanding: Effective for applications requiring analysis of combined text, image, and video data, such as document parsing, UI understanding, and video analysis.