Jackwang111/FaithEyes-8B-RL
FaithEyes-8B-RL is an 8 billion parameter vision-language model (VLM) developed by Haoqing Wang et al., initialized from Qwen3-VL-8B-Instruct and further enhanced with reinforcement learning (RL). It features a unique multi-agent self-judging framework where the VLM acts as both a main agent for visual question-answering with tool calls and a subagent that verifies the helpfulness of generated process images. This RL checkpoint specifically focuses on producing faithful tool calls by incorporating an accuracy-independent tool reward, making it particularly effective for tasks requiring reliable and relevant tool usage without external model dependencies.
Loading preview...
Overview
FaithEyes-8B-RL is an 8 billion parameter vision-language model (VLM) developed by Haoqing Wang et al. It is the final reinforcement learning (RL) checkpoint of the FaithEyes project, built upon the Qwen3-VL-8B-Instruct base model and fine-tuned from the FaithEyes-8B-SFT model. This model introduces a novel multi-agent self-judging framework designed to enhance the faithfulness of tool use in VLMs.
Key Capabilities
- Self-Judging Framework: The model operates as both a main agent, solving visual questions with code-based tool calls, and a subagent, which evaluates the helpfulness of process images generated by the main agent. This internal verification steers subsequent reasoning and scales tool rewards.
- Faithful Tool Use: Through a GRPO-based RL stage, FaithEyes-8B-RL is optimized to produce tool calls that genuinely capture the queried target, rather than merely decorative ones. It utilizes an accuracy-independent tool reward mechanism to maintain stable tool usage and prevent tool avoidance.
- Enhanced Performance: The RL stage significantly repairs performance drops in mathematical vision tasks introduced during SFT, pushing MathVista scores beyond both the SFT model and the base model.
- No External Dependencies: The self-judging subagent is instantiated by the model itself using a separate prompt, ensuring that the judgment process is available at inference without requiring any external models.
Good for
- Applications requiring reliable and contextually relevant tool execution in vision-language tasks.
- Scenarios where self-correction and internal verification of intermediate steps are crucial for accurate results.
- Tasks involving complex visual reasoning that benefit from guided tool interaction and robust mathematical understanding.