DavidAU/Qwen3.5-9B-The-Defiant-Fable-DARK-ROAST-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP

VISIONConcurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 7, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

DavidAU/Qwen3.5-9B-The-Defiant-Fable-DARK-ROAST-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP is a 9 billion parameter Qwen 3.5 family model, fine-tuned by DavidAU and Nightmedia for enhanced general intelligence and instruction following. This "DARK ROAST" edition features stronger de-censoring and improved performance, exceeding several benchmarks of larger Qwen 3.5 and 3.6 models in 4-bit and 8-bit quantization. It is designed for applications requiring high intelligence, superior instruction adherence, and uncensored responses, with a native context length of 262,144 tokens and vision capabilities.

Loading preview...

Model Overview

DavidAU/Qwen3.5-9B-The-Defiant-Fable-DARK-ROAST-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP is a 9 billion parameter model based on the Qwen 3.5 architecture, fine-tuned by DavidAU and Nightmedia. This "DARK ROAST" edition emphasizes stronger de-censoring while maintaining high performance. It is a "Heretic" model, meaning it is trained to follow user instructions without refusal, even for sensitive content.

Key Capabilities

  • Enhanced Intelligence & Instruction Following: Achieves superior general intelligence and instruction adherence through multi-stage fine-tuning and merging.
  • Benchmark Performance: Exceeds 7 critical benchmarks of the Qwen 3.5 27B model and, in some cases, meets Qwen 3.6 27B performance, even in 4-bit and 8-bit quantizations.
  • Uncensored Responses: Designed to generate content without refusals, offering greater flexibility for diverse applications.
  • Vision Capabilities: Supports image input, requiring a separate "mmproj" file for activation.
  • Multi-Token Prediction (MTP): Utilizes NEO IMATRIX GGUFs with MTP for potentially faster token generation, with specific settings recommended for optimal performance.
  • Long Context: Natively supports a context length of 262,144 tokens, extensible up to 1,010,000 tokens using YaRN scaling techniques.

Good For

  • Unrestricted Content Generation: Ideal for use cases where models typically refuse to generate certain types of content, such as creative writing or roleplay requiring explicit or sensitive themes.
  • High-Performance Applications on Smaller Hardware: Suitable for developers seeking strong performance from a 9B model, potentially outperforming larger models in specific benchmarks, especially with optimized 4-bit and 8-bit quantizations.
  • Complex Instruction Following: Excels in scenarios demanding precise and nuanced instruction adherence.
  • Multimodal Applications: Can be used for tasks involving image understanding when combined with the necessary vision encoder.