google/gemma-4-31B
Gemma 4 is a family of multimodal open models developed by Google DeepMind, capable of processing text, image, and audio inputs to generate text outputs. This release includes both pre-trained and instruction-tuned variants, featuring dense and Mixture-of-Experts (MoE) architectures with up to 30.7 billion parameters and a 256K token context window. Optimized for reasoning, coding, and agentic capabilities, Gemma 4 models support multilingual applications across over 140 languages and are deployable from mobile devices to servers.
Loading preview...
Gemma 4: Multimodal Models by Google DeepMind
Gemma 4 is a new family of open multimodal models from Google DeepMind, designed to handle text, image, and audio inputs (audio on smaller models) and generate text. These models come in various sizes, including E2B, E4B, 26B A4B (MoE), and 31B (Dense), making them suitable for deployment across a range of devices from high-end phones to servers. They feature a context window of up to 256K tokens and support over 140 languages.
Key Capabilities
- Multimodality: Processes text, images (with variable aspect ratio/resolution), video, and audio (E2B/E4B models).
- Reasoning: Includes configurable thinking modes for step-by-step problem-solving.
- Coding & Agentic Capabilities: Enhanced performance on coding benchmarks and native function-calling support.
- Diverse Architectures: Offers both Dense and Mixture-of-Experts (MoE) variants for scalable and efficient deployment.
- Long Context: Supports up to 256K tokens, utilizing a hybrid attention mechanism for efficiency.
- Native System Prompt Support: Enables more structured and controllable conversations.
Benchmark Highlights
Gemma 4 models demonstrate significant improvements over previous Gemma versions across various benchmarks. The 31B model achieves 85.2% on MMLU Pro, 89.2% on AIME 2026, and 80.0% on LiveCodeBench v6, showcasing strong performance in reasoning and coding. Multimodal benchmarks also show robust results, with the 31B model scoring 76.9% on MMMU Pro and 85.6% on MATH-Vision.
Good for
- Text Generation: Creative writing, chatbots, summarization.
- Multimodal Understanding: Image analysis, document parsing, video analysis, and audio processing (E2B/E4B).
- Coding: Code generation, completion, and correction.
- Agentic Workflows: Leveraging native function-calling for structured tool use.
- Research & Education: Foundation for VLM/NLP research and language learning tools.
Top 3 parameter combinations used by Featherless users for this model. Click a tab to see each config.