yah01/vjev-vision
yah01/vjev-vision is a 4.5 billion parameter listwise decision model with vision capabilities, built upon the Qwen3.5-4B base. It is designed to process combined text and image inputs, providing calibrated probabilities for typed questions (noul, choice, score) in a single forward pass without text generation. This model excels at tasks requiring visual understanding and structured decision-making, such as geometry questions and visual question answering, making it suitable for applications needing precise, probabilistic outputs from multimodal inputs.
Loading preview...
Overview
yah01/vjev-vision is a 4.5 billion parameter multimodal model, extending the Qwen3.5-4B base with vision capabilities. Unlike traditional LLMs that generate text, this model functions as a listwise decision model, providing calibrated probabilities for predefined question types: noul (true/false), choice (select one from options), and score (ordered scale). It processes both text and image inputs simultaneously, returning probabilistic answers in a single forward pass.
Key Capabilities
- Multimodal Input: Accepts text, images, or a combination of both as input states.
- Typed Questions: Supports specific question formats (
noul,choice,score) for structured decision-making. - Listwise Scoring: Options within a question compete, allowing for nuanced probabilistic outputs where options can influence each other.
- Vision Integration: Trained on geometry questions from COCO-2017 and VQAv2, enabling visual understanding tasks.
- Efficient Inference: Designed for direct probabilistic output without text generation, making it suitable for applications requiring calibrated confidence scores.
Performance Highlights
- Achieves 0.735 choice accuracy on COCO geometry questions and 0.706 choice accuracy on VQAv2, based on held-out data.
- Demonstrates strong performance on text-based tasks, maintaining 0.799 accuracy against human labels on held-out text data.
- Exhibits low rates of calling absent objects present (6.1% on POPE adversarial dataset), indicating robust object detection capabilities.
Good For
- Applications requiring probabilistic answers to structured questions from multimodal data.
- Visual question answering and scene understanding tasks.
- Decision support systems where calibrated confidence scores are crucial.
- Use cases needing to evaluate multiple options simultaneously, where options interact and influence each other's probabilities.