autotrust/JEV-27B-VL
autotrust/JEV-27B-VL is a 27 billion parameter multimodal large language model developed by autotrust, extending the JEV-27B model with vision capabilities. It is built upon the Qwen3.8-27B architecture and features a unique "System 1" decision-making component for rapid, calibrated probabilistic outputs on text and images, alongside a "System 2" for general text reasoning. This model excels at zero-shot visual decision tasks, such as robot arm control, computer interaction, video recommendation, and acting as an agent/multimodal judge, processing prompts up to 256K tokens.
Loading preview...
Overview
autotrust/JEV-27B-VL is a 27 billion parameter multimodal large language model developed by autotrust, extending the text-based JEV-27B with advanced vision capabilities. It integrates a unique "System 1" for rapid, calibrated probabilistic decisions on both text and images, and a "System 2" for general text reasoning, based on the Qwen3.8-27B architecture. This model is designed for high-speed, zero-shot decision-making in complex visual and textual environments, supporting context lengths up to 256K tokens.
Key Capabilities
- Rapid Visual Decision-Making (System 1): Provides calibrated probabilities for yes/no, rating (0-5), or choice (2-256 options) questions over images and text in a single forward pass. Achieves ~240 ms per decision for robot arm control and ~260 ms per click for computer use.
- Zero-Shot Visual Recommendation: Matches collaborative filtering performance for short-video recommendation using only cover images, solving the cold-start problem for new content.
- Advanced Agent and Multimodal Judging: Excels as an agent judge (73.2% accuracy on Plan-RewardBench, higher precision than all 16 leaderboard judges on AgentRewardBench) and a multimodal judge (78.3% accuracy on VL-RewardBench, competitive with GPT-5 on MMRB2 text-to-image tasks).
- Long Context Processing: Supports prompts up to 256K tokens, maintaining high accuracy for decisions dependent on single facts within very long documents.
- High Fidelity and Calibration: Demonstrates strong fidelity to JEV-27B's text decisions and high calibration (AUROC 0.995 for yes/no, ECE 0.0009).
Good For
- Automated Control Systems: Ideal for real-time decision-making in robotics (e.g., pick and place) and automated computer interaction (e.g., clicking UI elements).
- Content Recommendation: Particularly effective for cold-start scenarios in visual content recommendation, such as short-video feeds.
- AI Agent Evaluation: Serving as a robust judge for evaluating the performance and trajectories of other AI agents.
- Multimodal Content Moderation and Quality Assessment: Assessing the quality or appropriateness of multimodal content by judging images and associated text.
- Applications Requiring Fast, Calibrated Probabilistic Outputs: Any use case where quick, reliable, and quantifiable decisions are needed from complex visual and textual inputs.