DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP

VISIONPricing:Input $0.431 / Cached $0.0862 / Output $1.12Concurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 7, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP is a 9 billion parameter Qwen 3.5-based model developed by DavidAU and Nightmedia. This multi-stage fine-tuned and merged model significantly enhances general intelligence and instruction following, exceeding 7 critical benchmarks of the Qwen 3.5 27B model. It is fully uncensored, designed to follow instructions without refusal, and features a compacted, stronger reasoning block. The model is optimized for superior instruction following and general intelligence in a compact package, supporting a 256k context length and vision capabilities.

Loading preview...

Model Overview

DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP is a 9 billion parameter model built upon the Qwen 3.5 architecture, developed through a multi-stage fine-tuning and merging process by DavidAU and Nightmedia. This model prioritizes enhanced general intelligence and superior instruction following, demonstrating performance that exceeds 7 critical benchmarks of the larger Qwen 3.5 27B model, and in some cases, even matching Qwen 3.6 27B.

Key Capabilities and Features

  • Exceptional Performance: Achieves 0.649 ARC-C in bf16, outperforming base Qwen 3.5 9B, Qwen 3.5 27B, and Qwen 3.6 35B-A3B, and nearly matching Qwen 3.6 27B.
  • Uncensored & Heretic: Designed to follow user instructions without refusal, offering maximum flexibility.
  • Enhanced Reasoning: Features a compacted and significantly strengthened thinking/reasoning block.
  • Optimized Quantization: Utilizes NEO IMATRIX GGUFs for improved accuracy (2-4% over normal GGUFs) and long context performance, with the output tensor modified to 16-bit full precision.
  • Multi-Token Prediction (MTP): Includes MTP GGUFs for potential speed increases (up to 185 T/S on Q4_K_S) under specific settings (temp <= 1, rep pen = 1).
  • Vision Capable: Supports vision inputs, requiring a separate 'mmproj' file.
  • Extended Context: Offers a 256k context window.

Ideal Use Cases

  • Applications requiring a highly intelligent and compliant model for complex instruction following.
  • Scenarios where uncensored content generation is necessary.
  • Tasks benefiting from strong reasoning capabilities in a smaller parameter count.
  • Users seeking optimized performance with GGUF quantizations, including faster inference with MTP variants for suitable workloads.