PrimeIntellect/DeepSeek-R1-Distill-Qwen-7B

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 14, 2025Architecture:Transformer Featherless Exclusive Cold

PrimeIntellect/DeepSeek-R1-Distill-Qwen-7B is a 7.6 billion parameter language model. This model is a distilled version, likely leveraging the DeepSeek-R1 architecture and Qwen's capabilities. Its primary differentiator and specific use cases are not detailed in the provided information, suggesting it may be a foundational or general-purpose model.

Loading preview...

Model Overview

This model, PrimeIntellect/DeepSeek-R1-Distill-Qwen-7B, is a 7.6 billion parameter language model. The name suggests it is a distilled version, potentially combining elements from the DeepSeek-R1 architecture and the Qwen model family. As a distilled model, it aims to retain the performance characteristics of a larger model while being more efficient in terms of size and computational requirements.

Key Characteristics

  • Parameter Count: 7.6 billion parameters, offering a balance between capability and efficiency.
  • Context Length: Supports a substantial context window of 32768 tokens, allowing for processing and generating longer sequences of text.
  • Distillation: Implies a focus on optimized performance and resource usage, making it potentially suitable for deployment in environments with constrained resources.

Use Cases

Given the available information, this model is likely suitable for a range of general-purpose natural language processing tasks. Specific optimizations or unique capabilities are not detailed, but its parameter count and context length suggest it can handle tasks such as:

  • Text generation and completion
  • Summarization
  • Question answering
  • Chatbot development

Further details on its specific training data, evaluation benchmarks, and intended applications are not provided in the current model card.