Openintelligent123/Phi-4-mini-instruct

TEXT GENERATIONPricing:Input $0.32 / Cached $0.016 / Output $1.4Concurrent Unit Cost:1Model Size:3.8BQuant:BF16Context Size:32kPublished:Sep 2, 2026License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

Phi-4-mini-instruct is a 3.8 billion parameter instruction-tuned decoder-only Transformer model developed by Microsoft, part of the Phi-4 family. It features a 128K token context length and a 200K vocabulary, optimized for strong reasoning capabilities, particularly in math and logic. This model is designed for broad multilingual commercial and research use in memory/compute-constrained and latency-bound environments.

Loading preview...

Model Overview

Phi-4-mini-instruct is a 3.8 billion parameter instruction-tuned model from Microsoft's Phi-4 family, designed for efficiency and strong reasoning. It was built using synthetic data and filtered high-quality public websites, with a focus on reasoning-dense content. The model incorporates supervised fine-tuning and direct preference optimization for precise instruction adherence and robust safety. It features a 128K token context length and a 200K vocabulary, supporting broad multilingual applications.

Key Capabilities and Differentiators

  • Optimized for Reasoning: Excels in math and logic tasks, achieving 88.6% on GSM8K and 64.0% on MATH, outperforming many similarly sized models.
  • Efficiency: Designed for memory/compute-constrained environments and latency-bound scenarios, making it suitable for edge deployments.
  • Multilingual Support: Features a larger vocabulary (200K tokens) and improved multilingual capabilities, showing strong performance on benchmarks like MGSM (63.9%) and Multilingual MMLU (49.3%).
  • Enhanced Instruction Following: Benefits from advanced post-training techniques for better instruction adherence and function calling.
  • Context Length: Supports a substantial 128K token context window, allowing for processing longer inputs.

Ideal Use Cases

  • General Purpose AI Systems: Suitable as a building block for generative AI features requiring strong reasoning in a compact form.
  • Resource-Constrained Applications: Excellent for deployments where memory, compute, or latency are critical factors.
  • Research and Development: Accelerates research in language and multimodal models due to its high-quality data foundation and performance characteristics.