Jackwang111/FaithEyes-8B-RL

VISIONPricing:Input $0.727 / Output $5.405Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 29, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

FaithEyes-8B-RL is an 8 billion parameter vision-language model (VLM) developed by Haoqing Wang et al., initialized from Qwen3-VL-8B-Instruct and further enhanced with reinforcement learning (RL). It features a unique multi-agent self-judging framework where the VLM acts as both a main agent for visual question-answering with tool calls and a subagent that verifies the helpfulness of generated process images. This RL checkpoint specifically focuses on producing faithful tool calls by incorporating an accuracy-independent tool reward, making it particularly effective for tasks requiring reliable and relevant tool usage without external model dependencies.

Loading preview...

Overview

FaithEyes-8B-RL is an 8 billion parameter vision-language model (VLM) developed by Haoqing Wang et al. It is the final reinforcement learning (RL) checkpoint of the FaithEyes project, built upon the Qwen3-VL-8B-Instruct base model and fine-tuned from the FaithEyes-8B-SFT model. This model introduces a novel multi-agent self-judging framework designed to enhance the faithfulness of tool use in VLMs.

Key Capabilities

  • Self-Judging Framework: The model operates as both a main agent, solving visual questions with code-based tool calls, and a subagent, which evaluates the helpfulness of process images generated by the main agent. This internal verification steers subsequent reasoning and scales tool rewards.
  • Faithful Tool Use: Through a GRPO-based RL stage, FaithEyes-8B-RL is optimized to produce tool calls that genuinely capture the queried target, rather than merely decorative ones. It utilizes an accuracy-independent tool reward mechanism to maintain stable tool usage and prevent tool avoidance.
  • Enhanced Performance: The RL stage significantly repairs performance drops in mathematical vision tasks introduced during SFT, pushing MathVista scores beyond both the SFT model and the base model.
  • No External Dependencies: The self-judging subagent is instantiated by the model itself using a separate prompt, ensuring that the judgment process is available at inference without requiring any external models.

Good for

  • Applications requiring reliable and contextually relevant tool execution in vision-language tasks.
  • Scenarios where self-correction and internal verification of intermediate steps are crucial for accurate results.
  • Tasks involving complex visual reasoning that benefit from guided tool interaction and robust mathematical understanding.