SpatialAxiom/SpatialAxiom-9B

VISIONConcurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 2, 2026License:cc-by-nc-4.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

SpatialAxiom-9B is a 9 billion parameter open spatial intelligence model developed by D2I-ai, built on the Qwen3.5 VLM family. It excels in general spatial reasoning, including 3D relational inference, perspective taking, multi-view correspondence, and embodied video understanding. This model is trained with a spatial data-centric recipe and achieves leading results on various spatial benchmarks, making it ideal for applications requiring advanced spatial perception from images and videos.

Loading preview...

SpatialAxiom-9B: An Open Spatial Intelligence Model

SpatialAxiom-9B is a 9 billion parameter vision-language model (VLM) developed by D2I-ai, designed for advanced general spatial reasoning. Built upon the Qwen3.5 VLM architecture, it specializes in tasks such as 3D relational inference, perspective taking, multi-view correspondence, and embodied video understanding.

Key Capabilities and Features

  • Leading Spatial Reasoning: Achieves top performance on benchmarks like VSI-Bench, MMSI-Bench, MindCube, ViewSpatial, and EmbSpatial, often surpassing larger proprietary and open-source models.
  • Spatial Data-Centric Training: Utilizes a systematic taxonomy of spatial tasks, balanced task distribution, and data synthesis to enhance data quality, trained purely with full-parameter Supervised Fine-Tuning (SFT).
  • Qwen3.5 VLM Backbone: Inherits the robust Qwen3.5 architecture, maintaining a general-purpose multimodal design without task-specific architectural modifications.
  • Multimodal Input Support: Processes both image and video inputs for spatial analysis, with a native context length of 262,144 tokens, extensible up to 1,010,000 tokens.
  • Open-Weight Release: Fully compatible with transformers and vLLM for easy integration and deployment.

When to Use SpatialAxiom-9B

  • Applications requiring precise spatial understanding: Ideal for robotics, augmented reality, virtual reality, and autonomous navigation.
  • Research in spatial AI: Serves as a clean, strong baseline for further fine-tuning or reinforcement learning in spatial intelligence tasks.
  • Video analysis for spatial relationships: Excels at interpreting spatial layouts and movements within video content.

Note: SpatialAxiom-9B is fine-tuned for direct responses only and does not support 'thinking blocks' or enable_thinking mode during inference.