Scrappy-Doo/gemma-4-31B-it
Gemma 4 31B-it is a 30.7 billion parameter instruction-tuned multimodal language model developed by Google DeepMind, part of the Gemma 4 family. It processes text and image inputs, generating text outputs, and features a 256K token context window. This model excels in reasoning, coding, and agentic workflows, offering enhanced capabilities for complex tasks.
Loading preview...
The Gemma 4 family, developed by Google DeepMind, introduces multimodal models capable of handling text and image input (with audio support on smaller variants) and generating text output. This release includes both dense and Mixture-of-Experts (MoE) architectures, with models ranging from E2B to 31B parameters, designed for deployment across various devices.
Key Capabilities
- Multimodality: Processes Text, Image (with variable aspect ratio and resolution), and Video. E2B, E4B, and 12B models also support Audio.
- Reasoning: Features configurable thinking modes for step-by-step problem-solving.
- Extended Context Window: Supports up to 256K tokens for medium models (12B, 26B A4B, 31B) and 128K for smaller models (E2B, E4B).
- Enhanced Coding & Agentic Capabilities: Improved performance on coding benchmarks and native function-calling support for autonomous agents.
- Multilingual Support: Pre-trained on over 140 languages, with out-of-the-box support for 35+ languages.
- Native System Prompt Support: Enables more structured and controllable conversations.
What Makes This Model Different?
This 31B dense variant offers frontier-level performance for reasoning, agentic workflows, coding, and multimodal understanding. It employs a hybrid attention mechanism for efficient processing of long contexts and achieves strong benchmark results across MMLU Pro (85.2%), AIME 2026 (89.2%), and LiveCodeBench v6 (80.0%). The Gemma 4 models are rigorously evaluated for safety, showing significant improvements over previous Gemma versions.
Good for
- Complex Reasoning Tasks: Leveraging its built-in reasoning mode.
- Multimodal Applications: Integrating text, image, and video understanding.
- Code Generation and Agentic Workflows: Due to enhanced coding and function-calling capabilities.
- Long-Context Applications: Benefiting from the 256K token context window.