SpatialAxiom/SpatialAxiom-9B
SpatialAxiom-9B is a 9 billion parameter open spatial intelligence model developed by D2I-ai, built on the Qwen3.5 VLM family. It excels in general spatial reasoning, including 3D relational inference, perspective taking, multi-view correspondence, and embodied video understanding. This model is trained with a spatial data-centric recipe and achieves leading results on various spatial benchmarks, making it ideal for applications requiring advanced spatial perception from images and videos.
Loading preview...
SpatialAxiom-9B: An Open Spatial Intelligence Model
SpatialAxiom-9B is a 9 billion parameter vision-language model (VLM) developed by D2I-ai, designed for advanced general spatial reasoning. Built upon the Qwen3.5 VLM architecture, it specializes in tasks such as 3D relational inference, perspective taking, multi-view correspondence, and embodied video understanding.
Key Capabilities and Features
- Leading Spatial Reasoning: Achieves top performance on benchmarks like VSI-Bench, MMSI-Bench, MindCube, ViewSpatial, and EmbSpatial, often surpassing larger proprietary and open-source models.
- Spatial Data-Centric Training: Utilizes a systematic taxonomy of spatial tasks, balanced task distribution, and data synthesis to enhance data quality, trained purely with full-parameter Supervised Fine-Tuning (SFT).
- Qwen3.5 VLM Backbone: Inherits the robust Qwen3.5 architecture, maintaining a general-purpose multimodal design without task-specific architectural modifications.
- Multimodal Input Support: Processes both image and video inputs for spatial analysis, with a native context length of 262,144 tokens, extensible up to 1,010,000 tokens.
- Open-Weight Release: Fully compatible with
transformersandvLLMfor easy integration and deployment.
When to Use SpatialAxiom-9B
- Applications requiring precise spatial understanding: Ideal for robotics, augmented reality, virtual reality, and autonomous navigation.
- Research in spatial AI: Serves as a clean, strong baseline for further fine-tuning or reinforcement learning in spatial intelligence tasks.
- Video analysis for spatial relationships: Excels at interpreting spatial layouts and movements within video content.
Note: SpatialAxiom-9B is fine-tuned for direct responses only and does not support 'thinking blocks' or enable_thinking mode during inference.