Prototype3550/Qwen3.6-9B-Heretic-Uncensored-Thinking-Sweet-Madness

VISIONPricing:Input $1.6 / Cached $0.15 / Output $12Concurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 18, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

Prototype3550/Qwen3.6-9B-Heretic-Uncensored-Thinking-Sweet-Madness is a 27 billion parameter Qwen 3.6 model, uncensored via Heretic and then reduced to 9 billion parameters and 16 layers using Mergekit. Fine-tuned on 6 datasets over two stages, this model is designed for highly creative, uncensored, and unconventional text generation, excelling in imaginative and 'madness' style outputs. It features a 32768 token context length and retains image/video training capabilities from the full 27B model.

Loading preview...

Model Overview

Prototype3550/Qwen3.6-9B-Heretic-Uncensored-Thinking-Sweet-Madness is a unique 27 billion parameter Qwen 3.6 model that underwent a specialized process to become a 9 billion parameter, 16-layer variant. Initially uncensored using the Heretic method by P-E-W and 'trohrbaugh', it was then 'shrunk' via a modified Mergekit process. The resulting 9B model was subsequently fine-tuned on local hardware using Unsloth across 6 datasets in two stages, focusing on unifying its new layer structure.

Key Capabilities

  • Uncensored Generation: Designed to produce fully uncensored and unconventional text, pushing creative boundaries.
  • Creative Madness: Excels at generating highly imaginative, 'mad', and unique outputs, as demonstrated by its example generations.
  • Reduced Footprint: Achieves a 9B parameter count and 16 layers, making it faster (e.g., 200 t/s on Q4KS with a 5090 GPU) while retaining core capabilities.
  • Image/Video Training Intact: Lifts image and video training systems from the original 27B model.
  • High Context Length: Supports a substantial 32768 token context.

Good For

  • Creative Writing & Roleplay: Ideal for scenarios requiring highly imaginative, non-linear, or 'mad' narrative styles.
  • Unconventional Content Generation: Use cases where standard LLM guardrails are undesirable, and raw, unfiltered output is preferred.
  • Exploratory AI Research: For developers interested in pushing the boundaries of model behavior and exploring the effects of extreme fine-tuning and architectural modifications.

Important Considerations

  • Experimental Nature: Described as "Sweet Madness," this model is experimental and may require additional tuning for specific or even general use cases.
  • Performance vs. Qwen 3.5 9B: The README notes that Qwen 3.5 9B will outperform this model until it is fully tuned with 25-50k samples or specific datasets.
  • Resource Requirements: Can be tuned on Free Google Colab or local hardware with 12-16 GB VRAM.