autotrust/JEV-27B-VL

Hugging Face
VISIONPricing:Input $0.4 / Cached $0.15 / Output $3Concurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 30, 2026License:apache-2.0Architecture:Transformer0.3K Open Weights Featherless Exclusive Warm

autotrust/JEV-27B-VL is a 27 billion parameter multimodal large language model developed by autotrust, extending the JEV-27B model with vision capabilities. It is built upon the Qwen3.8-27B architecture and features a unique "System 1" decision-making component for rapid, calibrated probabilistic outputs on text and images, alongside a "System 2" for general text reasoning. This model excels at zero-shot visual decision tasks, such as robot arm control, computer interaction, video recommendation, and acting as an agent/multimodal judge, processing prompts up to 256K tokens.

Loading preview...

Overview

autotrust/JEV-27B-VL is a 27 billion parameter multimodal large language model developed by autotrust, extending the text-based JEV-27B with advanced vision capabilities. It integrates a unique "System 1" for rapid, calibrated probabilistic decisions on both text and images, and a "System 2" for general text reasoning, based on the Qwen3.8-27B architecture. This model is designed for high-speed, zero-shot decision-making in complex visual and textual environments, supporting context lengths up to 256K tokens.

Key Capabilities

  • Rapid Visual Decision-Making (System 1): Provides calibrated probabilities for yes/no, rating (0-5), or choice (2-256 options) questions over images and text in a single forward pass. Achieves ~240 ms per decision for robot arm control and ~260 ms per click for computer use.
  • Zero-Shot Visual Recommendation: Matches collaborative filtering performance for short-video recommendation using only cover images, solving the cold-start problem for new content.
  • Advanced Agent and Multimodal Judging: Excels as an agent judge (73.2% accuracy on Plan-RewardBench, higher precision than all 16 leaderboard judges on AgentRewardBench) and a multimodal judge (78.3% accuracy on VL-RewardBench, competitive with GPT-5 on MMRB2 text-to-image tasks).
  • Long Context Processing: Supports prompts up to 256K tokens, maintaining high accuracy for decisions dependent on single facts within very long documents.
  • High Fidelity and Calibration: Demonstrates strong fidelity to JEV-27B's text decisions and high calibration (AUROC 0.995 for yes/no, ECE 0.0009).

Good For

  • Automated Control Systems: Ideal for real-time decision-making in robotics (e.g., pick and place) and automated computer interaction (e.g., clicking UI elements).
  • Content Recommendation: Particularly effective for cold-start scenarios in visual content recommendation, such as short-video feeds.
  • AI Agent Evaluation: Serving as a robust judge for evaluating the performance and trajectories of other AI agents.
  • Multimodal Content Moderation and Quality Assessment: Assessing the quality or appropriateness of multimodal content by judging images and associated text.
  • Applications Requiring Fast, Calibrated Probabilistic Outputs: Any use case where quick, reliable, and quantifiable decisions are needed from complex visual and textual inputs.