Jackrong/DeepSeek-V4-Pro-Qwen3.5-4B
Jackrong/DeepSeek-V4-Pro-Qwen3.5-4B is a 4 billion parameter reasoning-focused language model, fine-tuned from Qwen3.5-4B. It is distilled from DeepSeek-V4-Pro in Max Effect mode, utilizing approximately 250,000 mathematics and STEM samples. This model excels at structured reasoning, mathematical problem-solving, and STEM question answering, offering compact deployment for local inference.
Loading preview...
What is DeepSeek-V4-Pro-Qwen3.5-4B?
DeepSeek-V4-Pro-Qwen3.5-4B is a 4 billion parameter language model developed by Jackrong, specifically fine-tuned for reasoning tasks. It is a distillation of the more powerful DeepSeek-V4-Pro (in Max Effect mode) into a smaller, more efficient Qwen3.5-4B base model. The training process involved approximately 250,000 mathematics and STEM samples, focusing on structured derivation, problem decomposition, and reliable answer construction.
Key Capabilities & Features
- Structured Reasoning: Learns multi-step solution patterns and explicit problem decomposition from a strong teacher model.
- Math & STEM Focus: Optimized for mathematical and scientific reasoning, trained on a large dataset of relevant samples.
- Compact Deployment: Designed for efficient local inference, bringing advanced reasoning capabilities to a 4B parameter class.
- Reproducible Evaluation: Benchmarked on complete GSM8K runs and fixed MMLU-Pro Math, Physics, and Chemistry samples.
Performance Highlights
The model achieves strong performance for its size, with a GSM8K score of 91.77% and an MMLU-Pro average of 76.47% across Math, Physics, and Chemistry. While it shows a performance gap compared to its 9B counterpart, it significantly outperforms DeepSeek-V3 on GSM8K, demonstrating effective knowledge transfer in a smaller footprint.
Should I use this for my use case?
This model is ideal for applications requiring strong grade-school and general mathematical problem-solving, lightweight physics, chemistry, and broader STEM question answering. It is particularly suited for local experiments with reasoning distillation and compact models, and for MTP-enabled llama.cpp deployment research. However, it has not been trained on coding data, and its performance on coding or tool-use tasks is not established. Users should verify high-stakes outputs due to its experimental nature and potential for arithmetic mistakes or reasoning failures on complex tasks.