RedHatAI/gemma-4-31B-it

VISIONConcurrent Unit Cost:2Model Size:31BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 18, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

RedHatAI/gemma-4-31B-it is an instruction-tuned multimodal language model from Google DeepMind, part of the Gemma 4 family. This 30.7 billion parameter dense model handles text and image inputs, generating text outputs, and features a 256K token context window. It excels in reasoning, coding, and agentic capabilities, making it suitable for complex multimodal understanding tasks.

Loading preview...

Gemma 4 31B-it: Multimodal Instruction-Tuned Model

RedHatAI/gemma-4-31B-it is an instruction-tuned model from Google DeepMind's Gemma 4 family, designed for multimodal understanding. This 30.7 billion parameter dense model supports text and image inputs, with a substantial 256K token context window. It is built for advanced reasoning, coding, and agentic workflows, featuring native function-calling support and enhanced coding benchmarks.

Key Capabilities

  • Multimodal Processing: Handles text and image inputs, with variable aspect ratio and resolution support for images. Video analysis is also supported by processing sequences of frames.
  • Reasoning: Incorporates a built-in reasoning mode that allows for step-by-step thinking before generating answers.
  • Extended Context: Features a 256K token context window, enabling deep awareness for long-context tasks.
  • Coding & Agentic Features: Achieves notable improvements in coding benchmarks and includes native function-calling for autonomous agents.
  • Multilingual Support: Pre-trained on over 140 languages, with out-of-the-box support for 35+ languages.
  • Native System Prompt Support: Introduces native support for the system role for more structured conversations.

Good For

  • Complex Reasoning Tasks: Its enhanced reasoning capabilities make it suitable for intricate problem-solving.
  • Code Generation and Completion: Excels in coding benchmarks, making it ideal for developer tools.
  • Multimodal Applications: Capable of processing and understanding both text and images, useful for diverse applications like document parsing, UI understanding, and image data extraction.
  • Agentic Workflows: Native function-calling support facilitates the development of highly capable autonomous agents.