rita-cohere/iolai-DeepSeek-R1-Distill-Qwen-7B

TEXT GENERATIONConcurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 16, 2026License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

The DeepSeek-R1-Distill-Qwen-7B model by DeepSeek-AI is a 7.6 billion parameter language model with a 32768 token context length, distilled from the larger DeepSeek-R1 reasoning model. It is based on the Qwen2.5 architecture and is specifically fine-tuned using reasoning data generated by DeepSeek-R1. This model demonstrates strong performance in reasoning, math, and code tasks, making it suitable for applications requiring robust analytical capabilities.

Loading preview...

Model Overview

DeepSeek-R1-Distill-Qwen-7B is a 7.6 billion parameter model developed by DeepSeek-AI, part of the DeepSeek-R1 series. This model is a distilled version of the larger DeepSeek-R1, which was trained using large-scale reinforcement learning (RL) to enhance reasoning capabilities without initial supervised fine-tuning (SFT). The distillation process transfers the reasoning patterns of the powerful DeepSeek-R1 into smaller, more efficient models like this Qwen-based variant.

Key Capabilities

  • Enhanced Reasoning: Inherits advanced reasoning patterns from DeepSeek-R1, which demonstrated self-verification, reflection, and long chain-of-thought generation.
  • Strong Performance: Achieves competitive results across various benchmarks, particularly in math (e.g., 92.8% on MATH-500 pass@1) and code (e.g., 1189 CodeForces rating).
  • Efficient Distillation: Proves that complex reasoning abilities can be effectively transferred from larger models to smaller, dense architectures, offering a powerful solution at a reduced scale.
  • Qwen2.5 Base: Built upon the Qwen2.5-Math-7B base model, leveraging its foundational strengths.

When to Use This Model

This model is particularly well-suited for:

  • Reasoning-intensive tasks: Ideal for applications requiring logical deduction, problem-solving, and complex analytical thinking.
  • Mathematical and Coding Challenges: Excels in benchmarks related to mathematics and code generation, making it a strong candidate for technical domains.
  • Resource-constrained environments: As a distilled model, it offers a balance of high performance and efficiency compared to its larger counterparts.
  • Research and Development: Provides a robust base for further fine-tuning or experimentation in reasoning-focused AI applications.