FriendliAI/Phi-3.5-vision-instruct
FriendliAI/Phi-3.5-vision-instruct is a 4.1 billion parameter multimodal model from the Phi-3 family, developed by Microsoft. It integrates an image encoder with the Phi-3 Mini language model, supporting a 128K token context length for both text and vision inputs. This model is optimized for general image understanding, optical character recognition, chart/table understanding, and multi-image reasoning, making it suitable for memory-constrained and latency-bound environments.
Loading preview...
FriendliAI/Phi-3.5-vision-instruct: A Compact Multimodal Powerhouse
FriendliAI/Phi-3.5-vision-instruct is a 4.1 billion parameter multimodal model developed by Microsoft, part of the Phi-3 family. It combines an image encoder with the Phi-3 Mini language model, offering a substantial 128K token context length for processing both text and visual information. The model is built on high-quality, reasoning-dense datasets, including synthetic data and filtered public web content, with a focus on both text and vision.
Key Capabilities
- Multimodal Understanding: Processes both text and images, excelling in general image understanding, optical character recognition (OCR), and chart/table interpretation.
- Multi-frame Reasoning: Enhanced to support multi-frame image understanding, enabling detailed image comparison, multi-image summarization, and video summarization.
- Optimized Performance: Designed for memory/compute-constrained environments and latency-bound scenarios, making it efficient for various applications.
- Robust Fine-tuning: Underwent rigorous supervised fine-tuning and direct preference optimization for precise instruction adherence and safety.
- Competitive Benchmarks: Shows improved performance on single-image benchmarks like MMMU (43.0) and MMBench (81.9), and competitive results on multi-frame benchmarks like BLINK (57.0 overall) and Video-MME (50.8 overall) against larger models.
Good For
- Developing general-purpose AI systems requiring visual and text input.
- Applications needing efficient image analysis in resource-limited settings.
- Research into language and multimodal models, serving as a foundational building block.
- Tasks involving multiple image comparisons or video clip summarization.