DreamFast/qwen3-8b-heretic

Hugging Face
TEXT GENERATIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Mar 20, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Warm

DreamFast/qwen3-8b-heretic is an abliterated version of Qwen's Qwen3-8B, processed using the Heretic v1.2.0 tool. This 8 billion parameter model is specifically modified to reduce refusals while maintaining quality, making it highly suitable as an uncensored text encoder for image generation models like Klein 9B. It is available in multiple quantization formats, including FP8, INT8, INT4, NVFP4, and MXFP8, all produced with SVD-guided learned rounding for enhanced fidelity.

Loading preview...

DreamFast/qwen3-8b-heretic: Abliterated Qwen3-8B

This model is an 'abliterated' version of the original Qwen/Qwen3-8B, created by DreamFast using the Heretic v1.2.0 tool. The primary goal of this modification is to significantly reduce model refusals while preserving the base model's quality, making it more versatile for certain applications.

Key Capabilities & Features

  • Reduced Refusals: Through 3000 optimization trials, the model's refusal rate was brought down from 100/100 to 13/100, meaning 87% of previously refused prompts now work.
  • Minimal Model Damage: A low KL Divergence of 0.0838 indicates that the abliteration process caused minimal degradation to the model's original capabilities.
  • Optimized for Image Generation: Specifically designed to function as an uncensored text encoder for image generation models, such as Klein 9B.
  • Extensive Quantization Support: Provided in numerous formats for various hardware and performance needs, including:
    • ComfyUI Formats: FP8, INT8 (ConvRot row-wise), INT4 (W4A4 ConvRot), NVFP4, and MXFP8. All use SVD-guided learned rounding for maximum fidelity.
    • GGUF Formats: F16, Q8_0, Q6_K, Q5_K_M, Q5_K_S, Q4_K_M (recommended balance), Q4_K_S, and Q3_K_M.
  • Hardware Optimization: NVFP4 and MXFP8 quantizations are optimized for Blackwell GPUs (RTX 5090/5080) for best performance, with software dequantization support for older GPUs.

Usage & Integration

This model can be easily integrated with ComfyUI, particularly for Klein 9B workflows, by placing the appropriate safetensors file in the ComfyUI/models/text_encoders/ directory. It also supports standard transformers library usage and llama.cpp via GGUF formats.

Popular Sampler Settings

Top 3 parameter combinations used by Featherless users for this model. Click a tab to see each config.

temperature
top_p
top_k
frequency_penalty
presence_penalty
repetition_penalty
min_p