InSight-doc/InSight-doc-8B

VISIONPricing:Input $0.727 / Output $5.405Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 4, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

InSight-doc/InSight-doc-8B is an 8 billion parameter vision-language agent developed by the authors of the InSight-doc paper, built upon Qwen3-VL-8B-Instruct. This model is specifically designed for long-document understanding, employing an agentic visual perception approach that uses a zoom-in tool for adaptive, high-resolution evidence acquisition. It excels at long-document visual question answering, significantly improving accuracy and reducing hallucination compared to its base model, while also offering substantial latency reductions.

Loading preview...

InSight-doc-8B: Agentic Visual Perception for Long-Document Understanding

InSight-doc-8B is an 8-billion parameter vision-language agent specialized in long-document understanding. Developed from Qwen3-VL-8B-Instruct, its core innovation lies in an agentic visual perception mechanism that adaptively uses visual resolution. The model starts with low-resolution page views and can call an image_zoom_in_tool to acquire high-resolution visual evidence from selected regions, enabling more precise answers.

Key Capabilities & Features

  • Agentic Visual Perception: Utilizes a coarse-to-fine approach, dynamically zooming into document regions for detailed analysis.
  • Enhanced Long-Document VQA: Specifically trained for visual question answering on lengthy documents.
  • Tool-Use Integration: Incorporates an image_zoom_in_tool for region-level visual evidence acquisition, crucial for its agentic behavior.
  • Improved Accuracy & Efficiency: Achieves 4.3-16.4 accuracy points improvement over Qwen3-VL-8B on benchmarks like DUDE, MP-DocVQA, MMLongBench-Doc, and LongDocURL.
  • Reduced Hallucination: Lowers hallucination rates by 40%+ on unanswerable questions and offers 1.7x-3.1x speedup (41%-68% latency reduction).
  • Two-Stage Training: Undergoes Supervised Fine-Tuning (SFT) on 17,913 zoom-in trajectories and Reinforcement Learning (RL) on 19,236 hard prompts.

When to Use This Model

InSight-doc-8B is ideal for applications requiring deep understanding and precise information extraction from long, visually complex documents. Its agentic approach makes it particularly effective for tasks where detailed visual evidence is critical, such as:

  • Document analysis and summarization
  • Complex visual question answering on documents
  • Automated data extraction from forms or reports

For optimal performance and to leverage its unique agentic capabilities, it is recommended to use InSight-doc-8B with its dedicated agent loop and image_zoom_in_tool rather than plain single-turn inference.