DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP
DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP is a 9 billion parameter Qwen 3.5-based causal language model developed by DavidAU and Nightmedia, fine-tuned for enhanced general intelligence and instruction following. It features a 32768 token context length and is designed to exceed the performance of larger Qwen 3.5 and 3.6 models on critical benchmarks, particularly in 4-bit and 8-bit quantization. This model is fully uncensored and includes vision capabilities, making it suitable for applications requiring robust, unconstrained responses and multimodal understanding.
Loading preview...
Model Overview
DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP is a 9 billion parameter model built on the Qwen 3.5 architecture, developed through a multi-stage fine-tuning and merging process by DavidAU and Nightmedia. It aims to significantly improve general intelligence and instruction following, even surpassing larger Qwen 3.5 27B and Qwen 3.6 27B/35B models in several critical benchmarks, particularly in 4-bit and 8-bit quantized versions. The model is notable for its "Heretic" training, meaning it is fully uncensored and designed to respond without refusal.
Key Capabilities
- Enhanced Intelligence & Instruction Following: Demonstrates superior performance in general intelligence and instruction adherence compared to its base model and some larger Qwen variants.
- Uncensored Output: Trained to provide direct responses without content refusals.
- Vision-Capable: Supports image inputs, requiring a separate
mmprojfile for activation. - Optimized Quantization: Achieves high performance (e.g., 640 ARC-C) in both 4-bit and 8-bit precision.
- NEO IMATRIX GGUFs: Utilizes NEO IMATRIX GGUFs for improved quantization accuracy (2-4% over standard GGUFs) and long context performance, with an output tensor modified to 16-bit full precision.
- Multi-Token Prediction (MTP): Offers specialized MTP GGUFs for faster inference (up to 185 T/S on Q4_K_S) under specific sampling parameters (temp <= 1, rep pen = 1).
- Extended Context: Natively supports a 262,144 token context length, extensible up to 1,010,000 tokens using YaRN scaling techniques.
Good for
- Applications requiring unconstrained content generation: Ideal for use cases where content filtering or refusals are undesirable.
- Resource-constrained environments: Its strong performance in 4-bit and 8-bit quantizations makes it efficient for deployment on consumer-grade hardware.
- Multimodal tasks: Suitable for applications integrating image understanding.
- High-performance inference: MTP GGUFs provide significant speed advantages for certain generation profiles.
- Complex reasoning and instruction-following tasks: Designed to excel in scenarios demanding high general intelligence and precise adherence to instructions.