google/gemma-4-31B

Hugging Face
VISIONConcurrent Unit Cost:2Model Size:31BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Mar 12, 2026License:apache-2.0Architecture:Transformer0.5K Open Weights Warm

Gemma 4 by Google DeepMind is a family of multimodal open models, including a 31B parameter variant, designed for text, image, and optionally audio input with text output. These models feature a context window up to 256K tokens and multilingual support across 140+ languages. Utilizing both Dense and Mixture-of-Experts architectures, Gemma 4 excels in reasoning, coding, and agentic workflows, with specific optimizations for on-device deployment in smaller variants.

Loading preview...

Overview

Gemma 4 is a family of multimodal open models developed by Google DeepMind, offering both pre-trained and instruction-tuned variants. These models are capable of processing text and image inputs (with E2B, E4B, and 12B models also supporting audio) and generating text outputs. Key features include a substantial context window of up to 256K tokens and broad multilingual support for over 140 languages.

Key Capabilities

  • Multimodality: Processes text, images (with variable aspect ratio and resolution), and video. E2B, E4B, and 12B models also natively support audio.
  • Reasoning: Designed with configurable thinking modes for enhanced reasoning capabilities.
  • Long Context: Supports context windows up to 128K tokens for smaller models and 256K tokens for medium models.
  • Coding & Agentic Workflows: Achieves significant improvements in coding benchmarks and includes native function-calling support for autonomous agents.
  • Diverse Architectures: Available in Dense (E2B, E4B, 12B, 31B) and Mixture-of-Experts (26B A4B) variants, optimized for various deployment scenarios from mobile to servers.
  • Native System Prompt Support: Facilitates more structured and controllable conversations.

Good For

  • Content Creation: Generating creative text, code, and marketing copy.
  • Conversational AI: Powering chatbots and virtual assistants.
  • Multimodal Understanding: Tasks involving image analysis, document parsing, video analysis, and audio processing (for supported models).
  • Research & Development: Serving as a foundation for VLM and NLP research, and developing agentic applications.

Popular Sampler Settings

Top 3 parameter combinations used by Featherless users for this model. Click a tab to see each config.

temperature
top_p
top_k
frequency_penalty
presence_penalty
repetition_penalty
min_p