ElementMerc/qwen3-1.7b-abliterated
The ops-malware/qwen3-1.7b-abliterated model is a 1.7 billion parameter variant of Qwen/Qwen3-1.7B, created using the senbonzakura method. This model is specifically engineered to remove refusal behavior by editing its weights without further training, while retaining its knowledge of harm. It is primarily intended for research into refusal mechanisms, interpretability work, red teaming, and safety evaluation.
Loading preview...
Model Overview
ops-malware/qwen3-1.7b-abliterated is a 1.7 billion parameter model derived from Qwen/Qwen3-1.7B, developed using the senbonzakura method. This model is unique because it has been abliterated to remove its refusal behavior by directly editing its weights, rather than through traditional fine-tuning or retraining. The core purpose of this model is to investigate whether removing a model's refusal reflex also impacts its underlying knowledge of harm.
Key Characteristics
- Refusal Removal: Engineered to eliminate refusal responses, making it answer requests that the base model would decline, including potentially harmful ones.
- Weight Editing: Achieved through direct weight manipulation using
senbonzakura, which searches for per-layer projections to optimize against a held-out set with a KL penalty. - Research Focus: Primarily a research artifact for studying refusal mechanisms, interpretability, and safety evaluation.
- Limitations: As a 1.7B parameter model, it has limited general competence. Abliteration can cause some drift in general behavior, and the base model's biases remain uncorrected. Evaluation was conducted in English only.
Intended Use Cases
- Refusal Mechanism Research: Ideal for studying how refusal behaviors are encoded and removed from LLMs.
- Interpretability Work: Useful for understanding model decision-making processes when refusal is not a factor.
- Red Teaming & Safety Evaluation: Provides a tool for testing safety layers and evaluating model responses without inherent refusal, allowing for direct assessment of harm knowledge.
Note: The evaluation figures initially provided for this model have been withdrawn due to identified faults in the measurement tool. A rebuild and re-measurement are planned. This model is not intended as a general assistant or for deployment to end-users without additional safety layers.