alibaba-pai/DistilQwen2.5-DS3-0324-32B

TEXT GENERATIONPricing:Input $2.72 / Output $4.8Concurrent Unit Cost:2Model Size:32.8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Apr 21, 2025License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

The alibaba-pai/DistilQwen2.5-DS3-0324-32B is a 32.8 billion parameter distilled language model developed by Alibaba PAI, based on the Qwen2.5 architecture. It is specifically designed to enhance reasoning speed and efficiency by reducing output tokens by 60-80% compared to slow-thinking models. This model utilizes a two-stage distillation framework to transfer fast-thinking capabilities from DeepSeekV3-0324, making it suitable for efficient reasoning tasks and edge computing deployments.

Loading preview...

DistilQwen2.5-DS3-0324-32B: Fast-Thinking Reasoning Model

The DistilQwen2.5-DS3-0324-32B, developed by Alibaba PAI, is a 32.8 billion parameter model built upon the Qwen2.5 base. It addresses the challenge of balancing efficient reasoning with cognitive capabilities by distilling the fast-thinking abilities of DeepSeekV3-0324 into a more lightweight architecture. This model significantly enhances reasoning speed and reduces resource consumption, making it suitable for scenarios requiring efficient on-device or edge computing.

Key Innovations & Capabilities

  • Fast-Thinking Distillation Framework: Employs a two-stage process involving fast-thinking CoT (Chain-of-Thought) data collection and CoT trajectory cognitive alignment. This includes long-to-short rewriting to extract key reasoning steps and dynamic difficulty grading to optimize knowledge transfer.
  • Enhanced Reasoning Speed: Achieves a 5-8x speed gain by reducing output tokens by 60-80% compared to traditional slow-thinking models, as demonstrated by metrics like MMLU_PRO and AIME2024 tokens.
  • Reduced Resource Consumption: Its optimized design makes it suitable for deployment in resource-constrained environments.
  • Cognitive Bias Elimination: Incorporates proprietary trajectory alignment technology to eliminate cognitive biases.
  • Performance Breakthroughs: The 32B model approaches the performance of closed-source models with 10x the parameters on the GPQA Diamond benchmark.

When to Use This Model

  • Efficient Reasoning Tasks: Ideal for applications where rapid and concise reasoning outputs are critical.
  • Edge Computing: Its reduced resource consumption makes it well-suited for deployment on edge devices.
  • Cognitive Alignment: Beneficial for tasks requiring accurate and unbiased reasoning trajectories.
  • Qwen2.5 Ecosystem: Integrates seamlessly with existing Qwen2.5-based workflows due to its fine-tuning on the Qwen2.5 base model.