RedHatAI/gemma-4-31B

VISIONConcurrent Unit Cost:2Model Size:31BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 15, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

RedHatAI/gemma-4-31B is a 31 billion parameter multimodal large language model developed by Google DeepMind, part of the Gemma 4 family. This model handles text and image inputs, generating text outputs, and features a 256K token context window. It is optimized for reasoning, coding, and agentic workflows, offering strong performance in complex tasks and multilingual support across over 140 languages.

Loading preview...

Gemma 4: Multimodal LLMs by Google DeepMind

Gemma 4 is a family of open multimodal models from Google DeepMind, designed to process text and image inputs (with audio support on smaller variants) and generate text outputs. This 31 billion parameter model features a substantial 256K token context window and supports over 140 languages. It incorporates both Dense and Mixture-of-Experts (MoE) architectures, making it suitable for diverse deployment scenarios from mobile devices to servers.

Key Capabilities

  • Multimodality: Processes text, images (with variable aspect ratio and resolution), and video. Smaller E2B/E4B models also support audio.
  • Reasoning: Designed with configurable thinking modes for enhanced reasoning capabilities.
  • Extended Context: Supports context windows up to 256K tokens.
  • Coding & Agentic Workflows: Achieves improved coding benchmarks and includes native function-calling for autonomous agents.
  • Native System Prompt Support: Enables more structured and controllable conversations.

Benchmark Highlights (Instruction-tuned 31B model)

  • MMLU Pro: 85.2%
  • AIME 2026 (no tools): 89.2%
  • LiveCodeBench v6: 80.0%
  • MMMU Pro (Vision): 76.9%

Intended Usage

  • Content Creation: Text generation, chatbots, summarization, image data extraction.
  • Research & Education: NLP/VLM research, language learning tools, knowledge exploration.

Limitations

  • Performance is influenced by training data biases and gaps.
  • May struggle with highly complex or open-ended tasks.
  • Can generate factually incorrect or outdated information.
  • May lack common sense reasoning in certain situations.