Iambackup/DeepSeek-R1-Distill-Qwen-14B

TEXT GENERATIONConcurrent Unit Cost:1Model Size:14.8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 7, 2026License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

Iambackup/DeepSeek-R1-Distill-Qwen-14B is a 14.8 billion parameter language model developed by DeepSeek-AI, distilled from the larger DeepSeek-R1 model and based on Qwen2.5-14B. It is specifically fine-tuned using reasoning data generated by DeepSeek-R1, demonstrating enhanced performance in mathematical, coding, and general reasoning tasks. This model offers a powerful, smaller alternative for applications requiring strong reasoning capabilities with a 32K context length.

Loading preview...

DeepSeek-R1-Distill-Qwen-14B: Reasoning Capabilities in a Smaller Package

This model is a 14.8 billion parameter distilled version of DeepSeek-R1, built upon the Qwen2.5-14B architecture. It leverages reasoning patterns from the larger DeepSeek-R1 model, which was developed using a novel large-scale reinforcement learning (RL) approach without initial supervised fine-tuning (SFT) to foster complex chain-of-thought (CoT) reasoning.

Key Capabilities

  • Enhanced Reasoning: Distilled from DeepSeek-R1, which demonstrated advanced reasoning behaviors like self-verification and reflection through RL.
  • Strong Performance: Achieves competitive results across math, code, and general reasoning benchmarks, outperforming many models in its size class.
  • Efficient Deployment: As a distilled model, it offers powerful reasoning capabilities in a more compact form factor compared to its larger counterparts, making it suitable for more resource-constrained environments.
  • Long Context: Supports a context length of 32,768 tokens, enabling processing of extensive inputs.

Good For

  • Mathematical Problem Solving: Excels in benchmarks like AIME 2024 and MATH-500, making it suitable for applications requiring strong mathematical reasoning.
  • Code Generation and Understanding: Demonstrates solid performance in coding tasks, including LiveCodeBench and Codeforces.
  • General Reasoning Applications: Ideal for tasks that benefit from step-by-step reasoning and complex problem-solving.
  • Research and Development: Provides a robust base for further research into model distillation and reasoning enhancement.