etri-vilab/MultiHopSpatial-Qwen3-VL-32B-Instruct

VISIONPricing:Input $0.416 / Output $1.664Concurrent Unit Cost:2Model Size:33.4BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 18, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

The etri-vilab/MultiHopSpatial-Qwen3-VL-32B-Instruct is a 33.4 billion parameter vision-language model developed by etri-vilab, based on Qwen3-VL-32B-Instruct. It is post-trained using Group Relative Policy Optimization (GRPO) on the MultihopSpatial-Train dataset. This model specializes in multi-hop compositional spatial reasoning tasks, enhancing its ability to understand complex spatial relationships in images. It is designed for applications requiring advanced visual understanding and inferential capabilities.

Loading preview...

Overview

The etri-vilab/MultiHopSpatial-Qwen3-VL-32B-Instruct is a 33.4 billion parameter vision-language model developed by etri-vilab. It is built upon the Qwen3-VL-32B-Instruct architecture and has been specifically post-trained using Group Relative Policy Optimization (GRPO). This specialized training utilizes the MultihopSpatial-Train dataset, which comprises 6,791 samples, to significantly enhance the model's capabilities in multi-hop compositional spatial reasoning.

Key Capabilities

  • Enhanced Multi-hop Spatial Reasoning: Excels at understanding and inferring complex spatial relationships that require multiple steps of reasoning.
  • Vision-Language Integration: Combines visual input processing with natural language understanding and generation, inherited from its Qwen3-VL base.
  • Instruction Following: Designed to follow instructions for visual question answering, including identifying objects and their spatial arrangements.
  • Bounding Box Prediction: Capable of providing bounding box coordinates for regions related to its answers, as demonstrated in the quick start example.

Good For

  • Complex Visual Question Answering: Ideal for scenarios where questions require inferring relationships between multiple objects or locations within an image.
  • Research in Spatial AI: A valuable tool for researchers exploring advanced spatial reasoning in vision-language models.
  • Applications Requiring Detailed Visual Analysis: Suitable for tasks demanding a deep understanding of visual scenes beyond simple object recognition.