sovitrath/Phi-3.5-vision-instruct
sovitrath/Phi-3.5-vision-instruct is a 4.1 billion parameter open multimodal model from Microsoft, part of the Phi-3 family, updated for the latest Transformers. It features a 128K token context length and excels in visual and text understanding, particularly for multi-frame image reasoning, optical character recognition, and chart/table understanding. This model is optimized for memory/compute constrained environments and latency-bound scenarios, offering enhanced performance on benchmarks like MMMU and MMBench.
Loading preview...
Model Overview
sovitrath/Phi-3.5-vision-instruct is an updated 4.1 billion parameter open multimodal model from the Phi-3 family, developed by Microsoft. It integrates an image encoder, connector, projector, and the Phi-3 Mini language model, supporting a substantial 128K token context length. The model is built on high-quality, reasoning-dense synthetic and filtered public data, with a strong focus on both text and vision.
Key Capabilities
- Multimodal Understanding: Processes both text and visual inputs, including general image understanding, OCR, and chart/table analysis.
- Multi-Frame Reasoning: Enhanced capabilities for detailed image comparison, multi-image summarization/storytelling, and video summarization, based on customer feedback.
- Performance Improvements: Shows improved performance on single-image benchmarks (e.g., MMMU from 40.2 to 43.0, MMBench from 80.5 to 81.9, TextVQA from 70.9 to 72.0).
- Optimized for Efficiency: Designed for memory/compute constrained environments and latency-bound scenarios.
- Rigorous Fine-tuning: Underwent supervised fine-tuning and direct preference optimization for instruction adherence and safety.
Good For
- Developing general-purpose AI systems requiring visual and text input.
- Applications needing efficient processing in resource-limited settings.
- Research and acceleration in language and multimodal models.
- Tasks involving detailed image comparison, multi-image summarization, and video summarization.