BAAI/Recon2Reason-Reasoning-4B

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 3, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

BAAI/Recon2Reason-Reasoning-4B is a 4.4 billion parameter vision-language model, fine-tuned from Qwen3-VL-4B, specifically designed for spatial reasoning in indoor and embodied scenes. It excels at understanding metric distance, relative position, object configuration, and spatial relations from visual inputs. This model retains the standard Qwen3-VL architecture and supports both single and multi-image visual inputs. It is optimized for tasks requiring precise spatial comprehension within visual environments.

Loading preview...

Recon2Reason Reasoning 4B Overview

Recon2Reason Reasoning 4B is a specialized 4.4 billion parameter vision-language model developed by BAAI. Fine-tuned from Qwen3-VL-4B, its primary focus is spatial reasoning within indoor and embodied environments, distinguishing it from general-purpose VLMs.

Key Capabilities and Features

  • Enhanced Spatial Reasoning: Demonstrates improved performance in understanding metric distance, relative object positioning, object configurations, and complex spatial relationships directly from visual inputs.
  • Qwen3-VL Architecture: Built upon the standard Qwen3VLForConditionalGeneration architecture, ensuring compatibility and ease of use without requiring custom code or trust_remote_code=True.
  • Visual Input Flexibility: Supports both single-image and multi-image visual inputs, crucial for comprehensive scene understanding.
  • Performance: Achieves strong results on benchmarks specifically designed for metric and qualitative spatial reasoning tasks.
  • Technical Specifications: Features 4,437,815,808 parameters, uses BF16 weights in sharded Safetensors format, and is tested with Transformers 4.57.1.

Ideal Use Cases

This model is particularly well-suited for applications requiring precise spatial understanding from visual data, such as:

  • Robotics: For navigation, object manipulation, and environmental interaction in indoor settings.
  • Augmented Reality/Virtual Reality: Enhancing spatial awareness and object placement within virtual environments.
  • Scene Understanding: Detailed analysis of indoor scenes for object relationships and layout.

It is important to note that this repository contains only the reasoning model; a separate release covers its retrieval-augmented scene-reconstruction extension.