alibaba-pai/DistilQwen2.5-DS3-0324-14B

TEXT GENERATIONPricing:Input $0.431 / Cached $0.0862 / Output $1.12Concurrent Unit Cost:1Model Size:14.8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Apr 21, 2025License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

The alibaba-pai/DistilQwen2.5-DS3-0324-14B is a 14.8 billion parameter model developed by Alibaba PAI, built on the Qwen2.5 base architecture. It is part of the DistilQwen2.5-DS3-0324 series, specifically designed for efficient reasoning by distilling fast-thinking capabilities from DeepSeekV3-0324. This model excels at reducing output tokens and resource consumption, making it suitable for applications requiring enhanced reasoning speed and edge computing deployment.

Loading preview...

Overview of DistilQwen2.5-DS3-0324-14B

The DistilQwen2.5-DS3-0324-14B model, developed by Alibaba PAI, is a 14.8 billion parameter language model based on the Qwen2.5 architecture. It is a key component of the DistilQwen2.5-DS3-0324 series, which focuses on transferring the efficient "fast-thinking" reasoning capabilities of DeepSeekV3-0324 to more lightweight models. This series addresses the challenge of balancing strong cognitive abilities with efficient reasoning.

Key Innovations and Capabilities

  • Fast-Thinking Distillation Framework: Utilizes a two-stage distillation process. The first stage involves collecting fast-thinking CoT (Chain-of-Thought) data through "Long-to-Short Rewriting" from DeepSeek-R1 and distilling rapid reasoning trajectories from DeepSeekV3-0324. The second stage focuses on CoT trajectory cognitive alignment using dynamic difficulty grading and an LLM-as-a-Judge validation mechanism.
  • Enhanced Reasoning Speed: Designed to significantly reduce output tokens (by 60-80% compared to slow-thinking models) and improve reasoning efficiency, as demonstrated by speed gains of 5-8x on benchmarks like MMLU_PRO and AIME2024 tokens for the 32B variant.
  • Reduced Resource Consumption: Its optimized design makes it suitable for deployment in resource-constrained environments, including edge computing.
  • Elimination of Cognitive Bias: Incorporates proprietary trajectory alignment technology to mitigate cognitive biases.
  • Open-Source Compatibility: Built upon the Qwen2.5 base model, ensuring compatibility and ease of use within the open-source ecosystem.

Ideal Use Cases

This model is particularly well-suited for applications where:

  • Efficient Reasoning is Critical: Tasks requiring quick, concise, and accurate reasoning outputs.
  • Resource Constraints Exist: Deployments on edge devices or environments with limited computational resources.
  • Cognitive Alignment is Important: Scenarios benefiting from reduced cognitive bias in reasoning processes.