etri-vilab/MultiHopSpatial-Qwen3-VL-4B-Instruct

VISIONConcurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Mar 20, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

The etri-vilab/MultiHopSpatial-Qwen3-VL-4B-Instruct is a 4 billion parameter vision-language model, based on Qwen3-VL-4B-Instruct, developed by etri-vilab. It is specifically post-trained using GRPO (Group Relative Policy Optimization) on the MultihopSpatial dataset to enhance multi-hop spatial reasoning capabilities. This model excels at understanding complex spatial relationships within images and answering related queries, making it suitable for advanced visual question answering tasks requiring sequential spatial inference.

Loading preview...

Overview

MultiHopSpatial-Qwen3-VL-4B-Instruct is a 4 billion parameter vision-language model developed by etri-vilab. It is built upon the Qwen3-VL-4B-Instruct base model and has been specifically fine-tuned for multi-hop spatial reasoning tasks. The model leverages GRPO (Group Relative Policy Optimization) during post-training on the dedicated MultihopSpatial dataset, which comprises 6,791 samples.

Key Capabilities

  • Enhanced Multi-hop Spatial Reasoning: Specialized training enables the model to process and infer complex spatial relationships that require multiple steps of reasoning.
  • Vision-Language Integration: Combines visual understanding with natural language processing to answer questions about spatial arrangements in images.
  • Qwen3-VL Architecture: Inherits the robust architecture of the Qwen3-VL series, providing a strong foundation for multimodal tasks.

Good For

  • Complex Visual Question Answering: Ideal for scenarios where understanding intricate spatial relationships between multiple objects in an image is crucial.
  • Research in Spatial AI: Useful for researchers exploring advanced spatial reasoning in vision-language models.
  • Applications Requiring Positional Inference: Can be applied in domains needing precise interpretation of object locations and relative positions based on visual input and textual queries.