Dmitriy-doc/DeepSeek-R1-Distill-Qwen-14B-bsl

TEXT GENERATIONConcurrent Unit Cost:1Model Size:14.8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 4, 2026Architecture:Transformer Featherless Exclusive Cold

Dmitriy-doc/DeepSeek-R1-Distill-Qwen-14B-bsl is a 14.8 billion parameter language model. This model is a distilled version, likely leveraging the DeepSeek-R1 architecture and Qwen base, designed for efficient performance. Its primary use case is general language understanding and generation tasks, offering a balance between size and capability.

Loading preview...

Model Overview

This model, Dmitriy-doc/DeepSeek-R1-Distill-Qwen-14B-bsl, is a 14.8 billion parameter language model. It is identified as a distilled version, suggesting an optimization process to achieve a more compact yet capable model, potentially drawing from the DeepSeek-R1 architecture and a Qwen base model. The model is designed for general-purpose language tasks.

Key Characteristics

  • Parameter Count: 14.8 billion parameters, indicating a substantial capacity for complex language understanding.
  • Context Length: Supports a context length of 32768 tokens, allowing it to process and generate longer sequences of text.
  • Distilled Model: Implies an efficiency-focused design, likely offering a good performance-to-resource ratio compared to larger, non-distilled counterparts.

Intended Use Cases

Given its parameter count and context window, this model is suitable for a variety of applications, including:

  • Text generation and completion.
  • Question answering.
  • Summarization.
  • Conversational AI.

Further details regarding specific training data, evaluation metrics, and fine-tuning procedures are not provided in the current model card.