liruiyangnb666/DeepSeek-R1-Distill-Qwen-7B

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 14, 2026License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

The liruiyangnb666/DeepSeek-R1-Distill-Qwen-7B is a 7.6 billion parameter language model distilled from DeepSeek-R1, developed by DeepSeek-AI. This model is fine-tuned using reasoning data generated by the larger DeepSeek-R1, which itself was developed through large-scale reinforcement learning to enhance reasoning capabilities. It is based on the Qwen2.5 architecture and is optimized for complex reasoning tasks across math, code, and general English and Chinese benchmarks, offering strong performance in a smaller, dense model format. The model supports a context length of 32768 tokens.

Loading preview...

Model Overview

DeepSeek-R1-Distill-Qwen-7B is a 7.6 billion parameter language model developed by DeepSeek-AI, part of a series of distilled models from the larger DeepSeek-R1. DeepSeek-R1 was initially trained using large-scale reinforcement learning (RL) to foster advanced reasoning behaviors, including self-verification and reflection, without initial supervised fine-tuning (SFT). The distillation process transfers these sophisticated reasoning patterns from the larger DeepSeek-R1 into smaller, more efficient models like this 7B variant, which is built upon the Qwen2.5 architecture.

Key Capabilities

  • Enhanced Reasoning: Inherits and distills advanced reasoning capabilities from DeepSeek-R1, which was specifically designed to excel in complex problem-solving.
  • Strong Performance: Demonstrates competitive performance across various benchmarks, including math (AIME 2024, MATH-500), code (LiveCodeBench, Codeforces), and general language tasks (MMLU, C-Eval).
  • Efficient Size: Provides powerful reasoning in a 7.6 billion parameter dense model, making it more accessible for deployment compared to much larger models.
  • Long Context: Supports a context length of 32768 tokens, enabling processing of extensive inputs.

When to Use This Model

  • Reasoning-Intensive Applications: Ideal for tasks requiring logical deduction, problem-solving, and multi-step reasoning in domains like mathematics and programming.
  • Resource-Constrained Environments: Suitable for scenarios where a smaller, yet highly capable, model is preferred over very large models, due to its distilled efficiency.
  • Research and Development: Useful for researchers exploring the transfer of reasoning capabilities through distillation and for developing applications that benefit from strong analytical skills.
  • Multilingual Tasks: Shows proficiency in both English and Chinese benchmarks, making it versatile for diverse language applications.