sovitrath/Phi-3.5-vision-instruct

VISIONPricing:Input $0.4 / Cached $0.02 / Output $0.8Concurrent Unit Cost:1Model Size:4.1BQuant:BF16Context Size:32kPublished:May 11, 2025License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

sovitrath/Phi-3.5-vision-instruct is a 4.1 billion parameter open multimodal model from Microsoft, part of the Phi-3 family, updated for the latest Transformers. It features a 128K token context length and excels in visual and text understanding, particularly for multi-frame image reasoning, optical character recognition, and chart/table understanding. This model is optimized for memory/compute constrained environments and latency-bound scenarios, offering enhanced performance on benchmarks like MMMU and MMBench.

Loading preview...

Model Overview

sovitrath/Phi-3.5-vision-instruct is an updated 4.1 billion parameter open multimodal model from the Phi-3 family, developed by Microsoft. It integrates an image encoder, connector, projector, and the Phi-3 Mini language model, supporting a substantial 128K token context length. The model is built on high-quality, reasoning-dense synthetic and filtered public data, with a strong focus on both text and vision.

Key Capabilities

  • Multimodal Understanding: Processes both text and visual inputs, including general image understanding, OCR, and chart/table analysis.
  • Multi-Frame Reasoning: Enhanced capabilities for detailed image comparison, multi-image summarization/storytelling, and video summarization, based on customer feedback.
  • Performance Improvements: Shows improved performance on single-image benchmarks (e.g., MMMU from 40.2 to 43.0, MMBench from 80.5 to 81.9, TextVQA from 70.9 to 72.0).
  • Optimized for Efficiency: Designed for memory/compute constrained environments and latency-bound scenarios.
  • Rigorous Fine-tuning: Underwent supervised fine-tuning and direct preference optimization for instruction adherence and safety.

Good For

  • Developing general-purpose AI systems requiring visual and text input.
  • Applications needing efficient processing in resource-limited settings.
  • Research and acceleration in language and multimodal models.
  • Tasks involving detailed image comparison, multi-image summarization, and video summarization.