DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP
DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP is a 9 billion parameter Qwen 3.5-based causal language model developed by DavidAU and Nightmedia, fine-tuned for extreme intelligence and superior instruction following. It achieves 640 ARC-C in both 8-bit and 4-bit quantization, exceeding 7 critical benchmarks of larger Qwen 3.5 27B and Qwen 3.6 27B models. This model is fully uncensored and features enhanced reasoning capabilities, making it suitable for tasks requiring high intelligence and direct responses.
Loading preview...
DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP
This 9 billion parameter model, developed by DavidAU and Nightmedia, is a multi-stage fine-tune and merge of several Qwen 3.5 models. Its primary goal was to significantly enhance general intelligence and instruction following, resulting in "jaw-dropping performance" in a compact package. The model is notable for exceeding 7 critical benchmarks of larger Qwen 3.5 27B and Qwen 3.6 27B models, even in 4-bit and 8-bit quantizations, achieving an ARC-C score of 640.
Key Capabilities
- Enhanced Intelligence & Instruction Following: Outperforms larger models in specific benchmarks, demonstrating superior understanding and execution of instructions.
- Uncensored Output: Designed as a "Heretic" model, it provides direct responses without refusal, though explicit direction may be needed for highly graphic or explicit content.
- Optimized Reasoning: Features a compacted and stronger thinking/reasoning block.
- Multimodal Support: Natively supports vision inputs and can process video, requiring a separate "mmproj" file for image functionality.
- High-Speed Inference: MTP (Multi-Token Prediction) GGUFs can achieve speeds exceeding 185 tokens/second on Q4_K_S quantizations, significantly faster than regular GGUFs.
- Extended Context Length: Supports a native context length of 262,144 tokens, extensible up to 1,010,000 tokens using YaRN scaling.
Good For
- Applications requiring a highly intelligent and performant model in a smaller parameter size.
- Use cases where uncensored and direct responses are preferred.
- Tasks benefiting from enhanced reasoning and instruction following, such as complex problem-solving or creative generation.
- Developers seeking fast inference speeds, especially with MTP GGUFs, for real-time applications.
- Multimodal applications involving image and video understanding.