Jackrong/DeepSeek-V4-Pro-Qwen3.5-4B

VISIONConcurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 24, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Jackrong/DeepSeek-V4-Pro-Qwen3.5-4B is a 4 billion parameter reasoning-focused language model, fine-tuned from Qwen3.5-4B. It is distilled from DeepSeek-V4-Pro in Max Effect mode, utilizing approximately 250,000 mathematics and STEM samples. This model excels at structured reasoning, mathematical problem-solving, and STEM question answering, offering compact deployment for local inference.

Loading preview...

What is DeepSeek-V4-Pro-Qwen3.5-4B?

DeepSeek-V4-Pro-Qwen3.5-4B is a 4 billion parameter language model developed by Jackrong, specifically fine-tuned for reasoning tasks. It is a distillation of the more powerful DeepSeek-V4-Pro (in Max Effect mode) into a smaller, more efficient Qwen3.5-4B base model. The training process involved approximately 250,000 mathematics and STEM samples, focusing on structured derivation, problem decomposition, and reliable answer construction.

Key Capabilities & Features

  • Structured Reasoning: Learns multi-step solution patterns and explicit problem decomposition from a strong teacher model.
  • Math & STEM Focus: Optimized for mathematical and scientific reasoning, trained on a large dataset of relevant samples.
  • Compact Deployment: Designed for efficient local inference, bringing advanced reasoning capabilities to a 4B parameter class.
  • Reproducible Evaluation: Benchmarked on complete GSM8K runs and fixed MMLU-Pro Math, Physics, and Chemistry samples.

Performance Highlights

The model achieves strong performance for its size, with a GSM8K score of 91.77% and an MMLU-Pro average of 76.47% across Math, Physics, and Chemistry. While it shows a performance gap compared to its 9B counterpart, it significantly outperforms DeepSeek-V3 on GSM8K, demonstrating effective knowledge transfer in a smaller footprint.

Should I use this for my use case?

This model is ideal for applications requiring strong grade-school and general mathematical problem-solving, lightweight physics, chemistry, and broader STEM question answering. It is particularly suited for local experiments with reasoning distillation and compact models, and for MTP-enabled llama.cpp deployment research. However, it has not been trained on coding data, and its performance on coding or tool-use tasks is not established. Users should verify high-stakes outputs due to its experimental nature and potential for arithmetic mistakes or reasoning failures on complex tasks.