inclusionAI/UI-Venus-2-9B

VISIONConcurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 26, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

UI-Venus-2-9B by inclusionAI is a 9 billion parameter general-purpose foundation GUI agent, initialized from Qwen3.5-9B, designed for unified operation across mobile applications, web platforms, and desktop operating systems. It features a closed-loop reasoning-action framework, expanded coverage across 170+ multilingual mobile apps, 50k+ websites, and 50+ professional desktop applications, and integrates safety-aware mechanisms. This model excels at GUI grounding, mobile, web, computer-use, and CAPTCHA benchmarks, offering robust real-world application capabilities.

Loading preview...

UI-Venus-2-9B: A Unified GUI Agent

UI-Venus-2-9B is a 9 billion parameter general-purpose foundation GUI agent developed by inclusionAI, initialized from Qwen3.5-9B. It is engineered to operate seamlessly across diverse environments, including mobile applications, web platforms, and desktop operating systems, utilizing a unified closed-loop reasoning-action framework. The model observes interfaces, reasons about task states, executes actions, and incorporates environmental feedback.

Key Capabilities & Differentiators

  • Broad Environmental Coverage: Extends to over 170 multilingual mobile apps (Chinese and English), 50,000+ live websites, and native desktop OS with 50+ professional applications.
  • Deep-Research Task Generation: Employs a sophisticated pipeline to ground generated instructions in actual application functionality, enhancing accuracy and executability.
  • Robust Verification: Utilizes trace-level and sample-level evaluators based on visual keypoints and multi-model voting for reliable reinforcement learning reward signals, robust against reward hacking.
  • Integrated Safety Mechanisms: Features safety-aware controls to manage consequential actions, significantly reducing attack success rates on benchmarks like OSBlind (from 90%+ to 18.5% for the 9B model).
  • Strong Benchmark Performance: Achieves near state-of-the-art results among comparable models across GUI grounding, mobile, web, computer-use, and CAPTCHA benchmarks.

Training Methodology

UI-Venus-2-9B is trained through a three-stage pipeline: Multimodal Mid-Training, Offline RL, and Multi-teacher On-policy Distillation. This process leverages a complementary mixture of five task families: Grounding, CAPTCHA, Mobile, Web, and Computer, with data collected at scale across various environments.