yuchenwu73/GeoBox-R1

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 28, 2026License:cc-by-nc-4.0Architecture:Transformer Open Weights Featherless Exclusive Cold

GeoBox-R1 by yuchenwu73 is a 4 billion parameter vision-language model based on Qwen3-VL-4B-Instruct, designed for unified remote sensing visual grounding. It excels at identifying objects in aerial and satellite imagery, generating either horizontal or oriented bounding boxes from a given image and referring expression. The model achieves strong performance with macro averages of 58.78% [email protected] for HBB and 47.32% [email protected] for OBB tasks, making it suitable for precise object localization in remote sensing applications.

Loading preview...

GeoBox-R1: Unified Remote Sensing Visual Grounding

GeoBox-R1 is a 4 billion parameter vision-language model developed by yuchenwu73, built upon the Qwen3-VL-4B-Instruct architecture. Its core capability is unified remote sensing visual grounding, meaning it can precisely locate objects in aerial and satellite images based on textual descriptions.

Key Capabilities & Training Innovations

  • Flexible Bounding Box Output: Given an image and a referring expression, GeoBox-R1 can generate either a horizontal bounding box (HBB) or an oriented bounding box (OBB) for the described object.
  • Two-Stage Training Methodology:
    • Curriculum-Guided SFT: The model is initially fine-tuned using a curriculum that progresses from HBB grounding to OBB grounding, and then to HBB-to-OBB chain-of-thought reasoning.
    • Geometric RL (GDPO): This second stage enhances geometric precision using rotated-IoU and adaptive Wasserstein rewards, notably without requiring a learned reward model. GDPO improves both OBB and HBB performance.

Performance Highlights

GeoBox-R1 demonstrates competitive performance, achieving the best macro averages among evaluated baselines for its parameter size. Key results include:

Use Cases

This model is particularly well-suited for applications requiring precise object detection and localization in remote sensing imagery, such as environmental monitoring, urban planning, and defense. Its ability to handle both HBB and OBB outputs provides versatility for various analytical needs.