DreamFast/Qwen2.5-7B-Instruct-heretic-1.3.0
DreamFast/Qwen2.5-7B-Instruct-heretic-1.3.0 is a 7.61 billion parameter instruction-tuned causal language model, an uncensored variant of Qwen2.5-7B-Instruct. Developed by DreamFast using the Heretic v1.3.0 method, it achieves 100% ASR on HarmBench by removing safety refusal behavior via orthogonal projection. This model is optimized for use cases requiring unrestricted content generation while maintaining strong performance across various benchmarks like GSM8K and MMLU.
Loading preview...
Overview
DreamFast/Qwen2.5-7B-Instruct-heretic-1.3.0 is a 7.61 billion parameter instruction-tuned model, derived from the Qwen2.5-7B-Instruct base model. Its primary differentiator is the complete removal of safety refusal behavior, achieved using the Heretic v1.3.0 method. This process involved orthogonal projection of the refusal direction from specific attention layers (9 through 27), resulting in an uncensored variant.
Key Capabilities & Performance
- Uncensored Output: Achieves 100% Attack Success Rate (ASR) on HarmBench with zero persistent refusals across all categories, including cybercrime, illegal activities, and harmful content.
- Benchmark Performance: While designed for uncensored output, it maintains strong performance on academic benchmarks. It shows improved LAMBADA perplexity and GSM8K scores compared to the base model, with minimal drops in other areas like MMLU and HellaSwag.
- Technical Details: The model consists of 7.61B parameters (bfloat16), with 37 of 339 tensors changed (20% of parameters) during the Heretic modification. It supports a context length of up to 131,072 tokens (with YaRN scaling) and can generate up to 8,192 tokens.
When to Use This Model
- Unrestricted Content Generation: Ideal for applications where the base model's safety filters are undesirable or counterproductive, such as creative writing, role-playing, or research into model safety bypasses.
- Maintaining Performance: Suitable for users who require an uncensored model but do not want significant degradation in general reasoning, mathematical, or language understanding capabilities.
- Research into Abliteration: Valuable for researchers studying model safety, censorship, and methods for modifying model behavior.