Jackrong/DeepSeek-V4-Pro-Qwen3.5-9B

VISIONConcurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 19, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Jackrong/DeepSeek-V4-Pro-Qwen3.5-9B is a 9 billion parameter reasoning-focused language model, fine-tuned from Qwen3.5-9B. It was distilled from DeepSeek-V4-Pro in Max Effect mode, with supervised training on approximately 250,000 mathematics and STEM samples. This model excels in structured reasoning, achieving 95.28% on GSM8K and 90.53% average on MMLU-Pro Math, Physics, and Chemistry, making it highly efficient for mathematical and scientific problem-solving tasks.

Loading preview...

DeepSeek-V4-Pro-Qwen3.5-9B Overview

This model is a 9 billion parameter language model, fine-tuned by Jackrong from the Qwen3.5-9B architecture. It is a distillation of responses generated by DeepSeek-V4-Pro in "Max Effect" mode, specifically optimized for reasoning tasks.

Key Capabilities & Differentiators

  • Reasoning Focus: Primarily trained on approximately 250,000 mathematics and STEM samples, emphasizing structured derivation and problem decomposition.
  • High Accuracy: Achieves 95.28% on GSM8K (4-run average) and 90.53% average on MMLU-Pro (Math, Physics, Chemistry), outperforming its base model and other 9B alternatives.
  • Inference Efficiency: Demonstrates significantly fewer tokens per correct answer compared to Qwen3.5-9B and Claude Mythos-distilled 9B on MMLU-Pro, offering 36.1% fewer tokens than Qwen3.5-9B.
  • Strong Instruction Following: Shows high compliance with output formats, returning single-letter answers in 97.2% of MMLU-Pro cases, a notable improvement over comparison models.
  • Cross-Domain Transfer: Exhibits small, preliminary gains in programming and tool-calling despite no direct coding supervision, suggesting improved general reasoning structure.

Recommended Use Cases

  • Mathematical and STEM Problem Solving: Ideal for tasks requiring step-by-step derivation and analytical reasoning in science and engineering.
  • Structured Reasoning: Suitable for applications demanding precise instruction following and structured output formats.
  • Research: Valuable for studying reasoning distillation and cross-domain transfer in supervised fine-tuning.
  • Experimental Tool-Calling: Can be explored for tool-calling or programming workflows where outputs are independently validated.