google/gemma-4-31B-it

Hugging Face
VISIONConcurrent Unit Cost:2Model Size:31BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Mar 11, 2026License:apache-2.0Architecture:Transformer3.5K Open Weights Warm

Gemma 4 31B-it is a 30.7 billion parameter instruction-tuned multimodal language model developed by Google DeepMind. Part of the Gemma 4 family, it handles text and image inputs, generating text outputs, and features a 256K token context window. This model excels at reasoning, coding, and agentic workflows, offering native function-calling support and configurable thinking modes.

Loading preview...

Overview

Gemma 4 is a family of open multimodal models from Google DeepMind, designed for text and image input with text output. The 31B-it variant is a 30.7 billion parameter instruction-tuned model featuring a 256K token context window and multilingual support across 140+ languages. It is built with a hybrid attention mechanism for efficient long-context processing and includes native system prompt support.

Key Capabilities

  • Multimodality: Processes text and images (with variable aspect ratio/resolution), and video (via frame sequences). The E2B, E4B, and 12B models also support audio.
  • Reasoning: Incorporates configurable thinking modes for step-by-step problem-solving.
  • Long Context: Supports context windows up to 256K tokens.
  • Coding & Agentic Workflows: Enhanced coding benchmarks and native function-calling for autonomous agents.
  • Efficient Architectures: Available in Dense and Mixture-of-Experts (MoE) variants, optimized for diverse deployment scenarios from mobile to servers.

Good for

  • Content Creation: Generating creative text, chatbots, and text summarization.
  • Research & Education: NLP and VLM research, language learning tools, and knowledge exploration.
  • Multimodal Understanding: Tasks involving image analysis, document parsing, video understanding, and (for specific models) audio processing like ASR and AST.

Popular Sampler Settings

Top 3 parameter combinations used by Featherless users for this model. Click a tab to see each config.

temperature
top_p
top_k
frequency_penalty
presence_penalty
repetition_penalty
min_p