Lexius/Phi-3.5-vision-instruct
Lexius/Phi-3.5-vision-instruct is a 4.1 billion parameter open multimodal model developed by Microsoft, part of the Phi-3 family. It supports a 32768 token context length and excels in general image understanding, optical character recognition, and multi-image reasoning. This model is optimized for memory/compute-constrained environments and latency-bound scenarios, offering robust visual and text input capabilities.
Loading preview...
Overview
Lexius/Phi-3.5-vision-instruct is a 4.1 billion parameter multimodal model from Microsoft's Phi-3 family, designed for both text and vision inputs. It features a 128K token context length and is built upon high-quality, reasoning-dense synthetic and filtered public datasets. The model has undergone rigorous supervised fine-tuning and direct preference optimization for instruction adherence and safety. A key enhancement in this release is improved multi-frame image understanding and reasoning, including detailed image comparison, multi-image summarization, and video summarization, with observed performance boosts on benchmarks like MMMU (43.0), MMBench (81.9), and TextVQA (72.0).
Key Capabilities
- Multimodal Input: Processes both text and visual data.
- Multi-frame Reasoning: Excels at understanding and summarizing multiple images or video clips.
- Image Understanding: Strong performance in general image understanding, OCR, and chart/table comprehension.
- Optimized Performance: Designed for memory/compute-constrained and latency-bound environments.
- Instruction Following: Enhanced through SFT and DPO for precise instruction adherence.
Good For
- Developing general-purpose AI systems requiring visual and text input.
- Applications needing efficient image understanding and multi-image comparison.
- Use cases in memory or compute-constrained environments.
- Accelerating research in language and multimodal models.