SpatialAxiom/SpatialAxiom-35B-A3B
SpatialAxiom-35B-A3B is a 35.1 billion parameter Mixture-of-Experts (MoE) vision-language model with 3 billion active parameters, developed by SpatialAxiom. Built on the Qwen3.5 VLM family, it is specifically designed for general spatial reasoning, excelling in 3D relational inference, perspective taking, multi-view correspondence, and embodied video understanding. The model achieves leading performance on various spatial benchmarks, making it suitable for applications requiring advanced spatial intelligence.
Loading preview...
SpatialAxiom-35B-A3B: An Open Spatial Intelligence Model
SpatialAxiom-35B-A3B is a 35.1 billion parameter Mixture-of-Experts (MoE) vision-language model, featuring 3 billion active parameters. Developed by SpatialAxiom, this model is built upon the Qwen3.5 VLM architecture, maintaining a general-purpose multimodal design without task-specific architectural modifications. Its strength is derived from a unique spatial data-centric training recipe, which includes a systematic taxonomy of spatial tasks, balanced task distribution, and high-quality data synthesis.
Key Capabilities
- Leading Spatial Reasoning: Achieves top performance on benchmarks such as VSI-Bench, MMSI-Bench, MindCube, ViewSpatial, and EmbSpatial, often surpassing larger proprietary and open-source models.
- 3D Relational Inference: Excels at understanding and inferring relationships in 3D spaces.
- Perspective Taking: Capable of interpreting scenes from different viewpoints.
- Multi-view Correspondence: Handles information from multiple visual perspectives.
- Embodied Video Understanding: Processes and understands spatial dynamics within video content.
- High Context Length: Supports a native context length of 262,144 tokens, extensible up to 1,010,000 tokens.
- Full-Parameter SFT: Trained purely with full-parameter Supervised Fine-Tuning, providing a clean base for further post-training or downstream fine-tuning.
Good For
- Applications requiring advanced spatial intelligence and reasoning.
- Research and development in 3D scene understanding, robotics, and embodied AI.
- Developers looking for a robust VLM backbone with strong spatial capabilities for further fine-tuning or reinforcement learning.
- Use cases involving image and video analysis where spatial relationships are critical, such as analyzing room layouts or understanding routes in videos. Note that it is fine-tuned for direct responses only and does not support thinking blocks or chain-of-thought modes.