SpatialAxiom/SpatialAxiom-35B-A3B

TEXT GENERATIONConcurrent Unit Cost:3Model Size:35.1BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 2, 2026License:cc-by-nc-4.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

SpatialAxiom-35B-A3B is a 35.1 billion parameter Mixture-of-Experts (MoE) vision-language model with 3 billion active parameters, developed by SpatialAxiom. Built on the Qwen3.5 VLM family, it is specifically designed for general spatial reasoning, excelling in 3D relational inference, perspective taking, multi-view correspondence, and embodied video understanding. The model achieves leading performance on various spatial benchmarks, making it suitable for applications requiring advanced spatial intelligence.

Loading preview...

SpatialAxiom-35B-A3B: An Open Spatial Intelligence Model

SpatialAxiom-35B-A3B is a 35.1 billion parameter Mixture-of-Experts (MoE) vision-language model, featuring 3 billion active parameters. Developed by SpatialAxiom, this model is built upon the Qwen3.5 VLM architecture, maintaining a general-purpose multimodal design without task-specific architectural modifications. Its strength is derived from a unique spatial data-centric training recipe, which includes a systematic taxonomy of spatial tasks, balanced task distribution, and high-quality data synthesis.

Key Capabilities

  • Leading Spatial Reasoning: Achieves top performance on benchmarks such as VSI-Bench, MMSI-Bench, MindCube, ViewSpatial, and EmbSpatial, often surpassing larger proprietary and open-source models.
  • 3D Relational Inference: Excels at understanding and inferring relationships in 3D spaces.
  • Perspective Taking: Capable of interpreting scenes from different viewpoints.
  • Multi-view Correspondence: Handles information from multiple visual perspectives.
  • Embodied Video Understanding: Processes and understands spatial dynamics within video content.
  • High Context Length: Supports a native context length of 262,144 tokens, extensible up to 1,010,000 tokens.
  • Full-Parameter SFT: Trained purely with full-parameter Supervised Fine-Tuning, providing a clean base for further post-training or downstream fine-tuning.

Good For

  • Applications requiring advanced spatial intelligence and reasoning.
  • Research and development in 3D scene understanding, robotics, and embodied AI.
  • Developers looking for a robust VLM backbone with strong spatial capabilities for further fine-tuning or reinforcement learning.
  • Use cases involving image and video analysis where spatial relationships are critical, such as analyzing room layouts or understanding routes in videos. Note that it is fine-tuned for direct responses only and does not support thinking blocks or chain-of-thought modes.