bbbbiia/DeepSeek-R1-Distill-Qwen-7B

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 12, 2026License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

The DeepSeek-R1-Distill-Qwen-7B model by DeepSeek-AI is a 7.6 billion parameter language model with a 32768 token context length, distilled from the larger DeepSeek-R1 reasoning model. It is fine-tuned using reasoning data generated by DeepSeek-R1, based on the Qwen2.5-Math-7B architecture. This model excels in mathematical, coding, and general reasoning tasks, demonstrating strong performance on benchmarks like AIME 2024 and MATH-500, making it suitable for applications requiring robust analytical capabilities.

Loading preview...

Model Overview

DeepSeek-R1-Distill-Qwen-7B is a 7.6 billion parameter language model developed by DeepSeek-AI, featuring a 32768 token context length. It is part of the DeepSeek-R1-Distill series, which focuses on transferring the advanced reasoning capabilities of the larger DeepSeek-R1 model into more compact, efficient architectures. This specific model is built upon the Qwen2.5-Math-7B base and has been fine-tuned using high-quality reasoning data generated by DeepSeek-R1.

Key Capabilities

  • Enhanced Reasoning: Benefits from distillation of DeepSeek-R1's reasoning patterns, which were developed through large-scale reinforcement learning (RL).
  • Strong Performance in Math & Code: Achieves competitive results on benchmarks such as AIME 2024 (55.5% pass@1), MATH-500 (92.8% pass@1), and LiveCodeBench (37.6% pass@1).
  • Efficient Size: Offers powerful reasoning in a 7.6B parameter model, making it more accessible for various deployment scenarios compared to much larger models.
  • Qwen2.5 Base: Leverages the robust architecture and pre-training of the Qwen2.5 series.

Ideal Use Cases

  • Mathematical Problem Solving: Excels in complex mathematical reasoning and problem-solving tasks.
  • Code Generation and Analysis: Suitable for applications requiring code understanding and generation.
  • General Reasoning Applications: Can be applied to a wide range of tasks demanding logical inference and structured thinking.
  • Resource-Constrained Environments: Provides strong performance for its size, making it a good choice where computational resources are a consideration.