FriendliAI/Phi-3.5-vision-instruct

VISIONConcurrent Unit Cost:1Model Size:4.1BQuant:BF16Context Size:32kPublished:Mar 4, 2025License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

FriendliAI/Phi-3.5-vision-instruct is a 4.1 billion parameter multimodal model from the Phi-3 family, developed by Microsoft. It integrates an image encoder with the Phi-3 Mini language model, supporting a 128K token context length for both text and vision inputs. This model is optimized for general image understanding, optical character recognition, chart/table understanding, and multi-image reasoning, making it suitable for memory-constrained and latency-bound environments.

Loading preview...

FriendliAI/Phi-3.5-vision-instruct: A Compact Multimodal Powerhouse

FriendliAI/Phi-3.5-vision-instruct is a 4.1 billion parameter multimodal model developed by Microsoft, part of the Phi-3 family. It combines an image encoder with the Phi-3 Mini language model, offering a substantial 128K token context length for processing both text and visual information. The model is built on high-quality, reasoning-dense datasets, including synthetic data and filtered public web content, with a focus on both text and vision.

Key Capabilities

  • Multimodal Understanding: Processes both text and images, excelling in general image understanding, optical character recognition (OCR), and chart/table interpretation.
  • Multi-frame Reasoning: Enhanced to support multi-frame image understanding, enabling detailed image comparison, multi-image summarization, and video summarization.
  • Optimized Performance: Designed for memory/compute-constrained environments and latency-bound scenarios, making it efficient for various applications.
  • Robust Fine-tuning: Underwent rigorous supervised fine-tuning and direct preference optimization for precise instruction adherence and safety.
  • Competitive Benchmarks: Shows improved performance on single-image benchmarks like MMMU (43.0) and MMBench (81.9), and competitive results on multi-frame benchmarks like BLINK (57.0 overall) and Video-MME (50.8 overall) against larger models.

Good For

  • Developing general-purpose AI systems requiring visual and text input.
  • Applications needing efficient image analysis in resource-limited settings.
  • Research into language and multimodal models, serving as a foundational building block.
  • Tasks involving multiple image comparisons or video clip summarization.