Jackrong/DeepSeek-V4-Pro-Qwen3.5-9B
VISIONConcurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 19, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold
Jackrong/DeepSeek-V4-Pro-Qwen3.5-9B is a 9 billion parameter reasoning-focused language model, fine-tuned from Qwen3.5-9B. It was distilled from DeepSeek-V4-Pro in Max Effect mode, with supervised training on approximately 250,000 mathematics and STEM samples. This model excels in structured reasoning, achieving 95.28% on GSM8K and 90.53% average on MMLU-Pro Math, Physics, and Chemistry, making it highly efficient for mathematical and scientific problem-solving tasks.
Loading preview...
DeepSeek-V4-Pro-Qwen3.5-9B Overview
This model is a 9 billion parameter language model, fine-tuned by Jackrong from the Qwen3.5-9B architecture. It is a distillation of responses generated by DeepSeek-V4-Pro in "Max Effect" mode, specifically optimized for reasoning tasks.
Key Capabilities & Differentiators
- Reasoning Focus: Primarily trained on approximately 250,000 mathematics and STEM samples, emphasizing structured derivation and problem decomposition.
- High Accuracy: Achieves 95.28% on GSM8K (4-run average) and 90.53% average on MMLU-Pro (Math, Physics, Chemistry), outperforming its base model and other 9B alternatives.
- Inference Efficiency: Demonstrates significantly fewer tokens per correct answer compared to Qwen3.5-9B and Claude Mythos-distilled 9B on MMLU-Pro, offering 36.1% fewer tokens than Qwen3.5-9B.
- Strong Instruction Following: Shows high compliance with output formats, returning single-letter answers in 97.2% of MMLU-Pro cases, a notable improvement over comparison models.
- Cross-Domain Transfer: Exhibits small, preliminary gains in programming and tool-calling despite no direct coding supervision, suggesting improved general reasoning structure.
Recommended Use Cases
- Mathematical and STEM Problem Solving: Ideal for tasks requiring step-by-step derivation and analytical reasoning in science and engineering.
- Structured Reasoning: Suitable for applications demanding precise instruction following and structured output formats.
- Research: Valuable for studying reasoning distillation and cross-domain transfer in supervised fine-tuning.
- Experimental Tool-Calling: Can be explored for tool-calling or programming workflows where outputs are independently validated.