clzoro/Qwen3.5-4B-KIMI-Distill
clzoro/Qwen3.5-4B-KIMI-Distill is a 4.5 billion parameter causal language model developed by Kassadin88, distilled from KIMI-K2.5 and fine-tuned from Qwen3.5-4B. It is specifically enhanced for reasoning tasks across domains like coding, science, and mathematics, trained on 554K high-quality reasoning traces totaling 2 billion tokens. This model also inherits vision-language capabilities from its base model and supports a 32K context length, making it suitable for complex problem-solving requiring detailed step-by-step logic.
Loading preview...
Qwen3.5-4B-KIMI-Distill: Reasoning-Enhanced Language Model
This model, developed by Kassadin88, is a 4.5 billion parameter language model fine-tuned from Qwen3.5-4B. Its core distinction lies in its reasoning enhancement, achieved through distillation from the powerful KIMI-K2.5 model. It was trained on an extensive dataset of 554,381 high-quality chain-of-thought reasoning samples, comprising approximately 2 billion tokens.
Key Capabilities
- Enhanced Reasoning: Specialized in generating detailed, step-by-step reasoning traces across various complex domains.
- Multi-Domain Expertise: Strong performance in coding (60% of training data), science (15%), mathematics (10%), computer science (5%), and logical reasoning (5%).
- Vision-Language Integration: Inherits multimodal capabilities from its Qwen3.5-4B base, allowing for processing of both text and visual inputs.
- Long Context Window: Supports a context length of 32,768 tokens, beneficial for intricate problems and extended dialogues.
Good For
- Complex Problem Solving: Ideal for tasks requiring logical deduction, mathematical calculations, and scientific explanations.
- Code Generation and Analysis: Excels in generating well-structured code and providing explanations for programming challenges.
- Educational Applications: Useful for generating detailed solutions and explanations for academic problems.
- Research in AI Reasoning: A valuable tool for exploring and developing advanced reasoning capabilities in smaller models.