yah01/vjev-vision-pilot

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 22, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

The yah01/vjev-vision-pilot is a 4.5 billion parameter listwise decision model with vision capabilities, built on a Qwen3.5-4B base. Developed by yah01, it processes text, images, or both, to answer typed questions (noul, choice, score) by returning calibrated probabilities for options in a single forward pass, without text generation. This pilot checkpoint, trained for 300 steps on vision, excels at spatial reasoning and VQA tasks, offering a unique approach to multimodal question answering.

Loading preview...

Overview

The yah01/vjev-vision-pilot is a 4.5 billion parameter listwise decision model with integrated vision capabilities. It is designed to take a state (comprising text, images, or both) and respond to specific typed questions—noul (true/false), choice (select one from options), and score (rate on an ordered scale). Unlike traditional LLMs, it returns calibrated probabilities for every option in a single forward pass, eliminating text generation. This model is a re-creation of the Jev API's functionality, enhanced with image processing.

Key Capabilities

  • Multimodal Input: Processes both text and images simultaneously for comprehensive understanding.
  • Listwise Decision Making: Evaluates multiple options for a question in a single forward pass, providing probability distributions.
  • Typed Questions: Supports noul, choice, and score question types, offering structured outputs.
  • Efficient Inference: Delivers answers without generating text, focusing on direct probability outputs.
  • Vision Training: This pilot checkpoint (300 steps) shows significant improvement in spatial reasoning and VQA tasks, achieving 0.665 choice accuracy on COCO geometry and 0.654 on VQAv2.

Good For

  • Structured Question Answering: Ideal for applications requiring precise, probability-based answers to predefined question types.
  • Multimodal Analysis: Suitable for scenarios where decisions depend on both visual and textual information.
  • Benchmarking & Research: As a pilot checkpoint, it's valuable for exploring listwise decision models and vision integration, particularly for those interested in the Jev API's approach.
  • Resource-Conscious Deployment: The model requires approximately 9 GB of memory in bf16/fp16, making it accessible for certain inference setups.