DreamFast/qwen3-8b-heretic

Hugging Face
TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Mar 20, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Warm

DreamFast/qwen3-8b-heretic is an abliterated version of Qwen's Qwen3-8B, created using the Heretic v1.2.0 tool. This 8 billion parameter model has significantly reduced refusals (13/100 vs 100/100 original) while maintaining model quality with a low KL divergence of 0.0838. It is primarily optimized as an uncensored text encoder for image generation models like Klein 9B, and is also suitable for general text generation tasks requiring less censorship.

Loading preview...

DreamFast/qwen3-8b-heretic: Abliterated Qwen3-8B

This model is an "abliterated" version of the Qwen3-8B base model, processed using the Heretic v1.2.0 tool. The primary goal of this abliteration was to reduce model refusals while preserving its original quality.

Key Characteristics & Performance

  • Reduced Refusals: Achieves a refusal rate of 13/100, a significant reduction from the original model's 100/100, meaning 87% of previously refused prompts now work.
  • Minimal Model Damage: Maintains a low KL Divergence of 0.0838, indicating that the abliteration process caused minimal degradation to the model's overall quality.
  • Optimized for Image Generation: Specifically designed to function as an uncensored text encoder for image generation models, such as Klein 9B.
  • Versatile Formats: Available in multiple formats for diverse use cases:
    • HuggingFace Format: For direct use with transformers library.
    • ComfyUI Format: Optimized for integration with ComfyUI, including bf16, fp8, and nvfp4 quantizations.
    • GGUF Format: For llama.cpp and ComfyUI-GGUF, with various quantizations like Q4_K_M (recommended balance).

Usage & Limitations

This model can be used with transformers, llama.cpp, and ComfyUI. While abliteration significantly reduces refusals, it does not eliminate them entirely. The model inherits all limitations of the base Qwen3-8B model. NVFP4 quantization offers the smallest file size and best performance on Blackwell GPUs, with software dequantization support for older GPUs.

Popular Sampler Settings

Top 3 parameter combinations used by Featherless users for this model. Click a tab to see each config.

temperature
top_p
top_k
frequency_penalty
presence_penalty
repetition_penalty
min_p