saidutta69/DeepSeek-R1-Distill-Qwen-14B-heretic
The DeepSeek-R1-Distill-Qwen-14B-heretic model by saidutta69 is a 14.8 billion parameter, 32K context length variant of DeepSeek-R1-Distill-Qwen-14B. It is specifically engineered to suppress refusal behaviors through targeted weight edits, rather than fine-tuning, preserving the base model's knowledge and instruction-following. This model is optimized for developers seeking a powerful reasoning model without built-in refusal guardrails, providing direct answers to prompts.
Loading preview...
Overview
This model, DeepSeek-R1-Distill-Qwen-14B-heretic, is a 14.8 billion parameter, 32K context length variant of the deepseek-ai/DeepSeek-R1-Distill-Qwen-14B model. It has been modified using the Heretic v1.4.0 tool, which employs directional ablation (abliteration) to remove refusal behaviors. Unlike traditional fine-tuning, abliteration directly edits specific weights responsible for refusal, ensuring the base model's core knowledge and instruction-following capabilities remain largely intact.
Key Capabilities
- Decensored Output: Suppresses refusal behaviors, providing direct answers to prompts that the base model might otherwise decline.
- Preserved Core Intelligence: Maintains the original DeepSeek-R1-Distill-Qwen-14B's reasoning and instruction-following abilities.
- Efficient Refusal Removal: Utilizes abliteration for targeted weight edits, avoiding the coherence degradation often seen with fine-tuning methods aimed at overriding refusals.
Good For
- Developers who require a powerful Qwen-distilled DeepSeek-R1 reasoning model without built-in refusal guardrails.
- Use cases demanding direct, unfiltered chain-of-thought responses.
- Deployment on hardware with 16-24 GB GPUs, or consumer hardware using GGUF quantizations (Q4_K_M, Q5_K_M, Q6_K, Q8_0).
Note: This model deliberately removes safety filtering. Users are responsible for its deployment and outputs.