SpatialAxiom/SpatialAxiom-9B

VISIONPricing:Input $0.431 / Cached $0.0862 / Output $1.12Concurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 2, 2026License:cc-by-nc-4.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

SpatialAxiom-9B is a 9 billion parameter open spatial intelligence model developed by D2I-ai, built on the Qwen3.5 VLM family. It is designed for general spatial reasoning, excelling in 3D relational inference, perspective taking, multi-view correspondence, and embodied video understanding. This compact dense model is trained with a spatial data-centric recipe and full-parameter SFT, offering a clean starting point for further post-training.

Loading preview...

SpatialAxiom-9B: Open Spatial Intelligence Model

SpatialAxiom-9B, developed by D2I-ai, is a 9 billion parameter vision-language model (VLM) built upon the Qwen3.5 VLM architecture. It is specifically designed for advanced spatial reasoning tasks, including 3D relational inference, perspective taking, multi-view correspondence, and embodied video understanding.

Key Capabilities & Features

  • Leading Spatial Reasoning Performance: Outperforms many proprietary and larger open-source models on benchmarks like VSI-Bench, MMSI-Bench, MindCube, ViewSpatial, and EmbSpatial.
  • Spatial Data-Centric Training: Utilizes a systematic taxonomy of spatial tasks, balanced task distribution, and data synthesis to enhance data quality, trained purely with full-parameter Supervised Fine-Tuning (SFT).
  • Qwen3.5 VLM Backbone: Inherits the general-purpose multimodal design of the Qwen3.5 VLM, ensuring broad applicability without task-specific architectural modifications.
  • High Context Length: Natively supports a context length of 262,144 tokens, extensible up to 1,010,000 tokens, crucial for complex spatial and video understanding tasks.
  • Open-Weight Release: Publicly available on Hugging Face, compatible with transformers and vLLM for easy integration and deployment.
  • Direct Response Optimization: Fine-tuned for direct responses, it does not support 'thinking blocks' or chain-of-thought modes, simplifying inference for specific use cases.

Good for

  • Applications requiring robust 3D relational inference and spatial understanding from images and videos.
  • Research and development in embodied AI and multi-view scene analysis.
  • Developers seeking a powerful, open-source VLM optimized for spatial intelligence tasks, with a clean SFT base for further fine-tuning.