Lexius/Phi-3.5-vision-instruct

VISIONPricing:Input $0.4 / Cached $0.02 / Output $0.8Concurrent Unit Cost:1Model Size:4.1BQuant:BF16Context Size:32kPublished:Jun 2, 2025License:mitArchitecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Lexius/Phi-3.5-vision-instruct is a 4.1 billion parameter open multimodal model developed by Microsoft, part of the Phi-3 family. It supports a 32768 token context length and excels in general image understanding, optical character recognition, and multi-image reasoning. This model is optimized for memory/compute-constrained environments and latency-bound scenarios, offering robust visual and text input capabilities.

Loading preview...

Overview

Lexius/Phi-3.5-vision-instruct is a 4.1 billion parameter multimodal model from Microsoft's Phi-3 family, designed for both text and vision inputs. It features a 128K token context length and is built upon high-quality, reasoning-dense synthetic and filtered public datasets. The model has undergone rigorous supervised fine-tuning and direct preference optimization for instruction adherence and safety. A key enhancement in this release is improved multi-frame image understanding and reasoning, including detailed image comparison, multi-image summarization, and video summarization, with observed performance boosts on benchmarks like MMMU (43.0), MMBench (81.9), and TextVQA (72.0).

Key Capabilities

  • Multimodal Input: Processes both text and visual data.
  • Multi-frame Reasoning: Excels at understanding and summarizing multiple images or video clips.
  • Image Understanding: Strong performance in general image understanding, OCR, and chart/table comprehension.
  • Optimized Performance: Designed for memory/compute-constrained and latency-bound environments.
  • Instruction Following: Enhanced through SFT and DPO for precise instruction adherence.

Good For

  • Developing general-purpose AI systems requiring visual and text input.
  • Applications needing efficient image understanding and multi-image comparison.
  • Use cases in memory or compute-constrained environments.
  • Accelerating research in language and multimodal models.