mason01kent/Qwen3.6-9B-Heretic-Uncensored-Thinking-Sweet-Madness

VISIONConcurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 7, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Qwen3.6-9B-Heretic-Uncensored-Thinking-Sweet-Madness by mason01kent is a 27 billion parameter Qwen 3.6 model, reduced to 9 billion parameters and 16 layers, then fine-tuned for creative and uncensored text generation. It excels at imaginative and unconventional outputs, including detailed 'thinking' processes. This model is optimized for creative writing and exploring non-standard responses, offering a unique, uncensored perspective.

Loading preview...

Model Overview

mason01kent's Qwen3.6-9B-Heretic-Uncensored-Thinking-Sweet-Madness is a unique 27 billion parameter Qwen 3.6 model that has undergone a significant transformation. It was initially uncensored using the Heretic method by P-E-W (performed by "trohrbaugh"), then "shrunk" to a 9 billion parameter model with only 16 layers (down from 64) using a modified Mergekit process. The resulting 9B model was subsequently fine-tuned on local hardware using Unsloth across 6 datasets in two stages, focusing on unifying its new layer structure.

Key Capabilities

  • Uncensored Output: Designed to be fully uncensored, allowing for unrestricted content generation.
  • Creative Text Generation: Excels at producing highly imaginative, unconventional, and "mad" creative content, often including detailed internal "thinking" processes.
  • Reduced Footprint: Despite its origins, the 9B model is efficient, achieving fast inference speeds (e.g., 200 t/s on a 5090 GPU) due to its reduced layer count.
  • Intact Image/Video Systems: The underlying image/video training and systems from the full 27B model are preserved.
  • High Context Length: Supports a context length of 256k tokens.

Important Considerations

  • Experimental Nature: Described as "Sweet Madness," this model is experimental and may require additional tuning for specific or general use cases.
  • Knowledge Gaps: Due to the unique compression process, some knowledge or skills from the original 27B model might be missing.
  • Performance vs. Qwen 3.5 9B: The developer notes that Qwen 3.5 9B may outperform this model until it is fully tuned with more samples.

Good For

  • Creative Writing & Roleplay: Ideal for generating highly imaginative, unconventional, and uncensored narratives.
  • Exploring Unrestricted AI Behavior: Useful for researchers or developers interested in the outputs of a truly uncensored model.
  • Experimentation & Fine-tuning: A strong base for further fine-tuning on specific creative or niche datasets, especially with tools like Unsloth on consumer hardware.