Openintelligent123/DeepSeek-R1-Distill-Qwen-14B

TEXT GENERATIONPricing:Input $0.431 / Cached $0.0862 / Output $1.12Concurrent Unit Cost:1Model Size:14.8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 2, 2026License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

DeepSeek-R1-Distill-Qwen-14B is a 14.8 billion parameter language model developed by DeepSeek AI, distilled from the larger DeepSeek-R1 reasoning model and based on the Qwen2.5-14B architecture. It is specifically fine-tuned using reasoning data generated by DeepSeek-R1 to enhance its performance on complex reasoning, math, and code tasks. This model offers strong reasoning capabilities in a smaller, more efficient package, making it suitable for applications requiring robust analytical problem-solving.

Loading preview...

DeepSeek-R1-Distill-Qwen-14B Overview

DeepSeek-R1-Distill-Qwen-14B is a 14.8 billion parameter model from DeepSeek AI, part of their DeepSeek-R1 series focused on advanced reasoning. This model is a distillation of the larger DeepSeek-R1, which was developed using a novel large-scale reinforcement learning (RL) approach without initial supervised fine-tuning (SFT) to foster emergent reasoning behaviors. The distillation process transfers the reasoning patterns of the powerful DeepSeek-R1 into smaller, more efficient models like this Qwen2.5-14B variant.

Key Capabilities & Features

  • Enhanced Reasoning: Benefits from reasoning data generated by the 671B parameter DeepSeek-R1, leading to strong performance in math, code, and general reasoning benchmarks.
  • Distilled Efficiency: Achieves high reasoning capabilities in a 14.8B parameter model, making it more accessible than its larger counterparts.
  • Qwen2.5 Base: Built upon the Qwen2.5-14B architecture, leveraging its established language understanding and generation.
  • Long Context: Supports a context length of 32,768 tokens, enabling processing of extensive inputs.

When to Use This Model

This model is particularly well-suited for:

  • Complex Problem Solving: Excels in tasks requiring multi-step reasoning, such as mathematical proofs, logical puzzles, and intricate coding challenges.
  • Code Generation & Analysis: Demonstrates strong performance in coding benchmarks, making it valuable for development-related applications.
  • Resource-Constrained Environments: Offers a powerful reasoning engine in a smaller parameter count compared to the original DeepSeek-R1, ideal for deployment where computational resources are a consideration.
  • Research & Development: Provides a robust base for further fine-tuning or experimentation in reasoning-focused AI applications.