kagakouko/Spatial-Interactor-Qwen3-VL-8B
The Spatial-Interactor Qwen3-VL-8B is an 8 billion parameter BF16 checkpoint developed by kagakouko, based on Qwen/Qwen3-VL-8B-Instruct. This model is specifically fine-tuned to learn spatial reasoning through interaction with the observable physical world. It integrates local world-state and ego-motion transitions using supervised fine-tuning and On-Policy Distillation (OPD) for long trajectories. The model excels at processing image/video and question inputs for advanced spatial understanding.
Loading preview...
Overview of Spatial-Interactor Qwen3-VL-8B
This model is an 8 billion parameter BF16 checkpoint, built upon the Qwen/Qwen3-VL-8B-Instruct architecture, developed by kagakouko. Its core innovation lies in its ability to learn spatial reasoning by interacting with the physical world, processing both image and video inputs. The model achieves this through a unique training methodology involving supervised fine-tuning (SFT) for local world-state and ego-motion transitions, followed by On-Policy Distillation (OPD) to integrate these transitions over extended trajectories.
Key Capabilities
- Advanced Spatial Reasoning: Designed to understand and reason about spatial relationships and interactions within visual data.
- Interactive Learning: Utilizes a training paradigm that simulates interaction with the observable physical world.
- On-Policy Distillation (OPD): Employs OPD to effectively integrate successive transitions, enabling reasoning over long interaction sequences.
- Standard Inference: At inference, the model operates like its base model, requiring no extra trace, reward model, or teacher branch, accepting standard image/video and question inputs.
Good for
- Applications requiring sophisticated visual spatial understanding.
- Research and development in interactive AI and embodied intelligence.
- Tasks involving ego-motion analysis and understanding dynamic environments from visual input.