NostraEmpire/mirror-deepseek-r1-distill-qwen-7b
NostraEmpire/mirror-deepseek-r1-distill-qwen-7b is a 7.6 billion parameter language model from DeepSeek-AI, distilled from the larger DeepSeek-R1 model and based on Qwen2.5-Math-7B. It is specifically fine-tuned to excel in reasoning tasks across math, code, and general problem-solving, leveraging reasoning patterns discovered through large-scale reinforcement learning. This model offers a 32K context length and demonstrates strong performance on benchmarks, making it suitable for applications requiring robust analytical capabilities.
Loading preview...
DeepSeek-R1-Distill-Qwen-7B: Reasoning-Optimized Language Model
This model is a 7.6 billion parameter distilled version of DeepSeek-R1, built upon the Qwen2.5-Math-7B architecture. Developed by DeepSeek-AI, it leverages advanced reasoning patterns derived from the larger DeepSeek-R1 model, which was trained using large-scale reinforcement learning (RL) without initial supervised fine-tuning (SFT) to foster complex chain-of-thought (CoT) capabilities.
Key Capabilities
- Enhanced Reasoning: Inherits and distills sophisticated reasoning abilities from DeepSeek-R1, excelling in mathematical, coding, and general reasoning tasks.
- Distillation Advantage: Demonstrates that reasoning patterns from larger, more complex models can be effectively transferred to smaller, dense models, achieving strong performance.
- Competitive Benchmarks: Shows strong results on various benchmarks, including AIME 2024, MATH-500, GPQA Diamond, and LiveCodeBench, often outperforming other models in its size class.
- Extended Context: Supports a context length of 32,768 tokens, allowing for processing longer inputs and generating more comprehensive responses.
Good For
- Complex Problem Solving: Ideal for applications requiring detailed step-by-step reasoning, such as mathematical proofs, code generation, and logical puzzles.
- Resource-Efficient Reasoning: Provides high-quality reasoning capabilities in a smaller, more deployable package compared to its larger DeepSeek-R1 counterpart.
- Research and Development: Useful for researchers exploring distillation techniques and the transfer of advanced reasoning skills to more compact models.
- Applications requiring robust analytical performance.