saidutta69/DeepSeek-R1-Distill-Qwen-1.5B-heretic

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Aug 1, 2026License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

The saidutta69/DeepSeek-R1-Distill-Qwen-1.5B-heretic is a 1.5 billion parameter Qwen2.5-class language model, derived from deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B. Developed by saidutta69 using the Heretic v1.4.0 abliteration method, this model has its refusal behaviors suppressed through targeted weight edits rather than fine-tuning. It retains the base model's distilled reasoning and instruction-following capabilities, making it suitable for developers seeking a small, efficient model for chain-of-thought style reasoning without refusal guardrails.

Loading preview...

What is DeepSeek-R1-Distill-Qwen-1.5B-heretic?

This model is a 1.5 billion parameter variant of the DeepSeek-R1-Distill-Qwen-1.5B, created by saidutta69. Its primary distinction is the removal of refusal behaviors using the "abliteration" technique (Heretic v1.4.0), which involves targeted weight edits to the attention output and MLP down-projections. This method aims to suppress refusals while preserving the base model's core knowledge, reasoning traces, and instruction-following capabilities, unlike traditional fine-tuning which can sometimes degrade coherence.

Key Characteristics:

  • Decensored: Refusal behaviors are suppressed, allowing the model to comply with requests the base model would typically refuse.
  • Efficient: A 1.5B parameter Qwen2.5-class model, designed to run on CPUs and consumer hardware, with quantized versions as small as ~1 GB.
  • Reasoning: Retains the distilled reasoning capabilities of the original DeepSeek-R1 model, suitable for chain-of-thought style reasoning.
  • Methodology: Utilizes "abliteration" for refusal suppression, a technique that directly edits specific weight directions responsible for refusal, leaving other network capabilities intact.

Should you use this model?

This model is intended for developers who require the distilled reasoning of DeepSeek-R1 without its inherent refusal behaviors. It is not a capability upgrade over the base model but rather a modification to its safety guardrails. Users should be aware that this model will comply with requests the base model would refuse, including potentially harmful ones, as there is no safety filtering layered on top. Responsible deployment is crucial.