DavidAU/Qwen3.5-9B-The-Defiant-Fable-DARK-ROAST-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP

VISIONPricing:Input $0.431 / Cached $0.0862 / Output $1.12Concurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 7, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

DavidAU/Qwen3.5-9B-The-Defiant-Fable-DARK-ROAST-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP is a fine-tuned Qwen 3.5 9B language model developed by DavidAU and Nightmedia, featuring enhanced uncensored capabilities and superior instruction following. This model excels in general intelligence and reasoning, outperforming several Qwen 3.5 27B and Qwen 3.6 27B benchmarks in a smaller 9B parameter package. It supports a 256k context length, vision capabilities, and offers optimized GGUF quants (NEO IMATRIX and MTP) for improved accuracy and speed, making it suitable for demanding, uncensored applications requiring high performance.

Loading preview...

Model Overview

DavidAU/Qwen3.5-9B-The-Defiant-Fable-DARK-ROAST-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP is a highly optimized, uncensored Qwen 3.5 9B model, fine-tuned by DavidAU and Nightmedia. It focuses on significantly boosting general intelligence and instruction following, achieving "absolute fire" performance with 640 ARC-C for both 8-bit and 4-bit quants. This model notably exceeds 7 critical benchmarks of the larger Qwen 3.5 27B model and, in some cases, matches Qwen 3.6 27B performance, despite its smaller size.

Key Capabilities

  • Enhanced Uncensored Output: "DARK ROAST" edition offers stronger de-censoring than the original Qwen 3.5, providing responses without refusal, even for sensitive content.
  • Superior Instruction Following & Reasoning: Designed to improve general intelligence and instruction adherence, with a compacted and stronger thinking/reasoning block.
  • Multimodal Support: Features activated vision capabilities, requiring a separate "mmproj" file for image processing, and supports video input (though not fully tested by the developer).
  • Optimized GGUF Quants: Utilizes NEO IMATRIX GGUFs for 2-4% improved accuracy and enhanced long context performance, alongside Multi-Token Prediction (MTP) GGUFs for potential speed increases (up to 185 T/S).
  • Extended Context Window: Supports a native context length of 256k tokens, extensible up to 1,010,000 tokens via YaRN scaling techniques.

Good for

  • Applications requiring unfiltered and direct responses without content moderation.
  • Tasks demanding high general intelligence and precise instruction following in a compact model size.
  • Use cases benefiting from multimodal input, including image and video understanding.
  • Developers seeking optimized local inference with advanced GGUF quantization for speed and accuracy.