Openintelligent123/Phi-4-reasoning

TEXT GENERATIONPricing:Input $0.28 / Output $0.56Concurrent Unit Cost:1Model Size:14.7BQuant:FP8Context Size:32kPublished:Sep 2, 2026License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

Phi-4-reasoning is a 14.7 billion parameter dense decoder-only Transformer model developed by Microsoft Research, fine-tuned from Phi-4. It is optimized for advanced reasoning tasks in math, science, and coding, utilizing supervised fine-tuning on chain-of-thought traces and reinforcement learning. With a 32k token context length, this model is designed to accelerate research in language models and serve as a building block for generative AI features in memory/compute-constrained and latency-bound environments, particularly for applications requiring strong reasoning and logic capabilities.

Loading preview...

Model Overview

Phi-4-reasoning is a 14.7 billion parameter dense decoder-only Transformer model developed by Microsoft Research. It is a fine-tuned version of the Phi-4 base model, specifically optimized for advanced reasoning tasks. The model was trained using supervised fine-tuning on a dataset of chain-of-thought traces and reinforcement learning, incorporating synthetic prompts and high-quality filtered data focused on math, science, and coding skills, as well as safety alignment data.

Key Capabilities & Features

  • Enhanced Reasoning: Designed to excel in complex reasoning, logic, and problem-solving across math, science, and coding domains.
  • Chain-of-Thought Output: Generates responses with a distinct reasoning chain-of-thought block followed by a summarization block.
  • Optimized for Efficiency: Intended for use in memory/compute-constrained and latency-bound environments.
  • Extensive Context Length: Supports a 32k token context length, allowing for more complex queries and longer chain-of-thought generation.
  • Robust Safety: Incorporates a comprehensive safety post-training approach via supervised fine-tuning, leveraging both open-source and in-house generated synthetic prompts adhering to Microsoft safety guidelines.

Performance Highlights

Phi-4-reasoning demonstrates strong performance on various reasoning benchmarks, often outperforming significantly larger open-weight models. Key evaluation areas include:

  • Reasoning Tasks: Evaluated on AIME, GPQA-Diamond, OmniMath, LiveCodeBench, 3SAT, TSP, BA Calendar, Maze, and SpatialMap.
  • General-Purpose Benchmarks: Assessed on Kitab (information retrieval), IFEval and ArenaHard (instruction following), HumanEvalPlus (code generation), and MMLU-Pro (multitask language understanding).
  • Comparative Performance: Achieves competitive scores against models like DeepSeek-R1-Distill-70B and approaches the performance of the full DeepSeek R1 model, despite its smaller size.

Intended Use Cases

This model is primarily designed to accelerate research in language models and serve as a building block for generative AI features, especially in scenarios requiring:

  • Memory/compute-constrained environments.
  • Latency-bound applications.
  • Strong reasoning and logic capabilities.

It is best suited for prompts in a chat format, and users should always use the specified ChatML template with a detailed system prompt for optimal inference.