DavidAU/Qwen3.5-9B-The-Defiant-Fable-DARK-ROAST-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP

VISIONPricing:Input $0.431 / Cached $0.0862 / Output $1.12Concurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 7, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

The DavidAU/Qwen3.5-9B-The-Defiant-Fable-DARK-ROAST-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP is a 9 billion parameter Qwen 3.5-based language model, fine-tuned by DavidAU and Nightmedia for enhanced general intelligence and instruction following. It features strong de-censoring capabilities and improved reasoning, exceeding several benchmarks of larger Qwen 3.5 and 3.6 models. This model is optimized for complex tasks requiring high intelligence and precise instruction adherence, including vision capabilities and multi-token prediction (MTP) for faster inference.

Loading preview...

Model Overview

DavidAU/Qwen3.5-9B-The-Defiant-Fable-DARK-ROAST-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP is a 9 billion parameter model built on the Qwen 3.5 architecture, developed through a multi-stage fine-tuning and merging process by DavidAU and Nightmedia. This "DARK ROAST" edition offers significantly stronger de-censoring while maintaining high performance. It is noted for its extreme intelligence and superior instruction following, with a compacted and strengthened thinking/reasoning block.

Key Capabilities

  • Enhanced Intelligence & Instruction Following: Achieves high scores across various benchmarks, exceeding several Qwen 3.5 27B and Qwen 3.6 35B-A3B metrics, particularly in ARC-C (640).
  • Uncensored & Heretic: Designed to follow user instructions without refusal, including sensitive content, by being trained post-"Heretic'ing."
  • Vision-Capable: Supports image inputs and can be extended to video understanding with an additional mmproj file.
  • Optimized Inference: Provides NEO IMATRIX GGUFs for improved quantization accuracy (2-4% over normal GGUFs) and MTP (Multi-Token Prediction) GGUFs for faster generation speeds (up to 185 T/S on Q4_K_S).
  • Long Context Support: Natively handles up to 262,144 tokens, extensible to 1,010,000 tokens using YaRN scaling techniques.

Good For

  • Applications requiring a highly intelligent and uncensored model for complex tasks.
  • Scenarios where precise instruction following is critical.
  • Use cases benefiting from vision capabilities, such as image analysis and multimodal interactions.
  • Developers seeking optimized inference speeds through MTP and enhanced quantization accuracy.