hongxingli/P2R-8B

VISIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 30, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

P2R-8B is an 8 billion parameter fine-grained visual reasoning model developed by Hongxing Li and collaborators, built upon Qwen3-VL-8B-Instruct. It operates under the Perceive-to-Reason (P2R) framework, which decouples perception from reasoning for enhanced visual understanding. This model excels in complex visual reasoning tasks, demonstrating significant performance improvements over its base model on benchmarks like V-Star and HR-Bench. It is particularly suited for applications requiring detailed analysis and logical inference from visual inputs.

Loading preview...

P2R-8B: Decoupling Perception and Reasoning for Visual Tasks

P2R-8B is an 8 billion parameter visual reasoning model developed by Hongxing Li and collaborators, designed for fine-grained visual reasoning. It is based on the powerful Qwen3-VL-8B-Instruct architecture and introduces a novel Perceive-to-Reason (P2R) framework.

Key Capabilities & Differentiators

  • Decoupled Perception and Reasoning: The P2R framework separates the perception stage from the reasoning stage, allowing for more robust and accurate visual understanding.
  • Enhanced Visual Reasoning Performance: P2R-8B significantly outperforms its base model, Qwen3-VL-Instruct-8B, across various visual reasoning benchmarks. For instance, it achieves 93.7 on V-Star (a +9.9 improvement), 81.5 on HR-Bench-4K (a +6.7 improvement), and 82.6 on HR-Bench-8K (a +12.5 improvement).
  • Role-Aware Alternating RL Training: The model's training utilizes PRA-GRPO, a specialized role-aware alternating Reinforcement Learning strategy, contributing to its superior performance.

Good For

  • Applications requiring advanced visual reasoning and understanding.
  • Tasks that benefit from a decoupled approach to perception and reasoning.
  • Research and development in fine-grained visual analysis and complex image interpretation.