google/gemma-4-31B-it-qat-q4_0-unquantized

Hugging Face
VISIONConcurrent Unit Cost:2Model Size:31BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Apr 28, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Warm

The google/gemma-4-31B-it-qat-q4_0-unquantized model is a 30.7 billion parameter instruction-tuned multimodal language model developed by Google DeepMind. Part of the Gemma 4 family, this unquantized QAT checkpoint is optimized for reasoning, coding, and agentic workflows, supporting text and image inputs with a 256K token context window. It is designed for high-end consumer GPUs and servers, offering advanced capabilities in multimodal understanding and generation.

Loading preview...

Overview

This model is a 30.7 billion parameter instruction-tuned variant from the Gemma 4 family, developed by Google DeepMind. It is an unquantized Quantization-Aware Training (QAT) checkpoint, meaning it preserves high quality while being optimized for reduced memory footprint. Gemma 4 models are multimodal, processing text and image inputs (with some variants also supporting audio) and generating text outputs. This specific model features a 256K token context window and is designed for deployment on consumer GPUs and servers.

Key Capabilities

  • Multimodal Understanding: Processes text and image inputs, with variable aspect ratio and resolution support. Video understanding is also supported by processing sequences of frames.
  • Reasoning: Designed as a highly capable reasoner with configurable thinking modes.
  • Coding & Agentic Capabilities: Achieves notable improvements in coding benchmarks and includes native function-calling support for autonomous agents.
  • Long Context: Supports a substantial context window of up to 256K tokens.
  • Multilingual Support: Pre-trained on over 140 languages and offers out-of-the-box support for 35+ languages.
  • Native System Prompt Support: Introduces native support for the system role for more structured conversations.

Good For

  • Complex Reasoning Tasks: Its design and configurable thinking modes make it suitable for tasks requiring deep reasoning.
  • Code Generation and Agentic Workflows: Enhanced coding capabilities and function-calling support are ideal for development and agent-based applications.
  • Multimodal Applications: Excels in scenarios requiring the interpretation of both text and images, such as document parsing, visual question answering, and content creation involving visual elements.
  • High-Performance Deployments: Optimized for high-end hardware, making it suitable for demanding applications on consumer GPUs and servers.