sheldonrobinson/Phi-3.5-vision-instruct

VISIONPricing:Input $0.4 / Cached $0.02 / Output $0.8Concurrent Unit Cost:1Model Size:4.1BQuant:BF16Context Size:32kPublished:Oct 25, 2024License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

Microsoft's Phi-3.5-vision-instruct is a 4.1 billion parameter multimodal model from the Phi-3 family, designed for efficient visual and text understanding. It supports a 128K token context length and excels in memory/compute-constrained environments and latency-bound scenarios. This model is particularly strong in general image understanding, OCR, chart/table analysis, and multi-image/video summarization, offering enhanced multi-frame reasoning capabilities.

Loading preview...

Model Overview

Microsoft's Phi-3.5-vision-instruct is a lightweight, state-of-the-art open multimodal model with 4.1 billion parameters, part of the Phi-3 family. It integrates an image encoder, connector, projector, and the Phi-3 Mini language model, supporting a substantial 128K token context length. The model has undergone rigorous enhancement through supervised fine-tuning and direct preference optimization for precise instruction adherence and safety.

Key Capabilities

  • Multimodal Understanding: Processes both text and visual inputs, including single and multi-frame images.
  • Efficient Performance: Optimized for memory/compute-constrained environments and latency-bound scenarios.
  • Advanced Visual Reasoning: Excels in general image understanding, optical character recognition (OCR), chart and table understanding, and multi-image/video clip summarization.
  • Enhanced Multi-frame Reasoning: Features improved capabilities for detailed image comparison, multi-image summarization/storytelling, and video summarization, based on customer feedback.
  • Strong Benchmarks: Shows competitive performance on benchmarks like MMMU, MMBench, TextVQA, and particularly strong results in multi-frame benchmarks like BLINK and Video-MME, often outperforming models of similar size.

Good for

  • Applications requiring visual and text input capabilities in resource-limited settings.
  • Tasks involving general image understanding and analysis.
  • Use cases needing OCR, chart/table interpretation, or comparison of multiple images.
  • Developing generative AI features that leverage multimodal data, especially for summarization of image sequences or video clips.