DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP

VISIONConcurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 7, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP is a 9 billion parameter Qwen 3.5-based causal language model developed by DavidAU and Nightmedia, fine-tuned for enhanced general intelligence and instruction following. It features a 32768 token context length and is designed to exceed the performance of larger Qwen 3.5 and 3.6 models on critical benchmarks, particularly in 4-bit and 8-bit quantization. This model is fully uncensored and includes vision capabilities, making it suitable for applications requiring robust, unconstrained responses and multimodal understanding.

Loading preview...

Model Overview

DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP is a 9 billion parameter model built on the Qwen 3.5 architecture, developed through a multi-stage fine-tuning and merging process by DavidAU and Nightmedia. It aims to significantly improve general intelligence and instruction following, even surpassing larger Qwen 3.5 27B and Qwen 3.6 27B/35B models in several critical benchmarks, particularly in 4-bit and 8-bit quantized versions. The model is notable for its "Heretic" training, meaning it is fully uncensored and designed to respond without refusal.

Key Capabilities

  • Enhanced Intelligence & Instruction Following: Demonstrates superior performance in general intelligence and instruction adherence compared to its base model and some larger Qwen variants.
  • Uncensored Output: Trained to provide direct responses without content refusals.
  • Vision-Capable: Supports image inputs, requiring a separate mmproj file for activation.
  • Optimized Quantization: Achieves high performance (e.g., 640 ARC-C) in both 4-bit and 8-bit precision.
  • NEO IMATRIX GGUFs: Utilizes NEO IMATRIX GGUFs for improved quantization accuracy (2-4% over standard GGUFs) and long context performance, with an output tensor modified to 16-bit full precision.
  • Multi-Token Prediction (MTP): Offers specialized MTP GGUFs for faster inference (up to 185 T/S on Q4_K_S) under specific sampling parameters (temp <= 1, rep pen = 1).
  • Extended Context: Natively supports a 262,144 token context length, extensible up to 1,010,000 tokens using YaRN scaling techniques.

Good for

  • Applications requiring unconstrained content generation: Ideal for use cases where content filtering or refusals are undesirable.
  • Resource-constrained environments: Its strong performance in 4-bit and 8-bit quantizations makes it efficient for deployment on consumer-grade hardware.
  • Multimodal tasks: Suitable for applications integrating image understanding.
  • High-performance inference: MTP GGUFs provide significant speed advantages for certain generation profiles.
  • Complex reasoning and instruction-following tasks: Designed to excel in scenarios demanding high general intelligence and precise adherence to instructions.