OmniJev/OneJev-9B

VISIONPricing:Input $0.431 / Cached $0.0862 / Output $1.12Concurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 27, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

OmniJev/OneJev-9B is a 9 billion parameter multimodal System One decision model, fine-tuned from Qwen/Qwen3.5-9B. Developed by OmniJev, it processes screenshots, photos, videos, or text inputs to return calibrated probabilities for decision options in a single forward pass. This model is optimized for rapid, agent-based decision-making tasks, offering high accuracy on unseen questions and fast inference speeds.

Loading preview...

OneJev-9B: A Multimodal System One Decision Model

OneJev-9B, developed by OmniJev, is a 9 billion parameter multimodal model designed for System One decision-making. It is a full fine-tune of Qwen/Qwen3.5-9B, trained on 99,193 questions derived from real agent runs, videos, and images. This model excels at processing diverse inputs, including screenshots, photos, videos, and plain text, to provide calibrated probabilities for various decision options in a single pass.

Key Capabilities

  • Multimodal Input Processing: Accepts images (screenshots, photos), video, and text as input.
  • System One Decision Making: Returns calibrated probabilities for decision options, suitable for automated agent tasks.
  • High Accuracy: Demonstrates strong performance on a test set of unseen questions, outperforming text-only models like Jev 1.13.
  • Fast Inference: Achieves rapid response times, answering 1 question in 81 ms and 10 questions in 131 ms (13.1 ms/question) on an H200 GPU with a 1280x720 screenshot.
  • API Compatibility: Supports TypeSafe's System One API with an added media field for visual inputs.
  • Flexible Deployment: Can be run with PyTorch or via llama.cpp using GGUF builds, with image support in both and video support in the PyTorch server.

Good For

  • Automated agents requiring rapid, multimodal decision-making.
  • Applications needing to interpret visual and textual context for calibrated probabilistic outputs.
  • Scenarios where quick inference on diverse inputs (screenshots, video frames) is critical.