Lord1337iuu/DeepSeek-R1-Distill-Qwen-1.5B

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 22, 2026License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

DeepSeek-R1-Distill-Qwen-1.5B is a 1.5 billion parameter causal language model developed by DeepSeek AI, distilled from the larger DeepSeek-R1 model. It is fine-tuned on reasoning data generated by DeepSeek-R1, leveraging a Qwen2.5-Math-1.5B base. This model excels in mathematical, code, and general reasoning tasks, demonstrating that complex reasoning patterns can be effectively transferred to smaller models.

Loading preview...

DeepSeek-R1-Distill-Qwen-1.5B Overview

This model is a 1.5 billion parameter language model developed by DeepSeek AI, part of the DeepSeek-R1-Distill series. It is built upon the Qwen2.5-Math-1.5B base model and has been fine-tuned using reasoning data generated by the larger DeepSeek-R1 model. The core innovation lies in demonstrating that complex reasoning capabilities can be effectively distilled from powerful, larger models into significantly smaller, more efficient ones.

Key Capabilities

  • Reasoning Performance: Achieves strong performance across mathematical, code, and general reasoning benchmarks, benefiting from the distillation process.
  • Efficiency: As a 1.5B parameter model, it offers a more efficient solution for deploying reasoning-focused applications compared to larger counterparts.
  • Distilled Intelligence: Leverages advanced reasoning patterns learned by DeepSeek-R1, which was developed using large-scale reinforcement learning without initial supervised fine-tuning.

When to Use This Model

  • Resource-Constrained Environments: Ideal for applications requiring strong reasoning abilities where computational resources are limited.
  • Mathematical and Code Tasks: Particularly well-suited for problems involving mathematical reasoning and code generation, as indicated by its Qwen2.5-Math base.
  • Research and Development: Useful for exploring the efficacy of knowledge distillation techniques for transferring complex reasoning skills to smaller models.