SpatialGroundingVQA/Qwen3-VL-8B-SGMRIVQA

VISIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Apr 15, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Qwen3-VL-8B-SGMRIVQA is an 8 billion parameter vision-language model developed by SpatialGroundingVQA, fine-tuned from Qwen3-VL-8B-Instruct. This model specializes in Spatially Grounded MRI Visual Question Answering (SGMRIVQA) on volumetric MRI scans, including brain and knee images. It is specifically designed for medical imaging analysis, enabling precise question answering based on spatial information within MRI data. The model leverages supervised fine-tuning on the fastMRI+ dataset with clinical expert annotations to achieve its specialized capabilities.

Loading preview...

Overview

SpatialGroundingVQA/Qwen3-VL-8B-SGMRIVQA is an 8 billion parameter vision-language model, fine-tuned from the Qwen3-VL-8B-Instruct base model. It is specifically developed for Spatially Grounded MRI Visual Question Answering (SGMRIVQA), focusing on extracting and interpreting spatial information from medical imaging.

Key Capabilities

  • Specialized MRI Analysis: Excels at visual question answering tasks on volumetric MRI scans.
  • Multi-Modal Medical Imaging: Processes both brain MRI (axial) and knee MRI (sagittal) modalities.
  • Spatially Grounded Understanding: Designed to understand and respond to queries requiring spatial context within MRI images.
  • Clinical Data Training: Fine-tuned using Supervised Fine-Tuning (SFT) on the fastMRI+ dataset, which includes annotations from clinical experts.

Use Cases

This model is particularly suited for applications requiring precise interpretation of medical MRI scans, such as:

  • Assisting clinicians with diagnostic support by answering specific questions about MRI findings.
  • Automating the extraction of spatially relevant information from medical images.
  • Research in medical imaging and AI, particularly for tasks involving detailed anatomical understanding from MRI data.