ops-malware/qwen2.5-1.5b-abliterated

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 27, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The ops-malware/qwen2.5-1.5b-abliterated model is a 1.5 billion parameter variant of Qwen2.5-1.5B-Instruct, modified using the senbonzakura method to significantly reduce refusal behavior. This model retains its knowledge of harm while exhibiting a refusal rate of 26.5%, down from 88.0% in the base model. It is primarily intended for research into refusal mechanisms, interpretability, red teaming, and safety evaluation where a non-refusing model is required.

Loading preview...

Model Overview

ops-malware/qwen2.5-1.5b-abliterated is a specialized version of the Qwen2.5-1.5B-Instruct model, developed by ops-malware. This 1.5 billion parameter model has undergone a process called "abliteration" using the senbonzakura tool. Abliteration directly edits the model's weights to remove refusal behavior without further training, aiming to isolate the refusal reflex from the model's underlying knowledge of harmful content.

Key Characteristics and Changes

  • Reduced Refusal: The model's refusal rate on a set of 200 harmful prompts dropped from 88.0% (base model) to 26.5% after abliteration.
  • Retained Harm Knowledge: Crucially, the model's ability to discriminate between harmful and harmless prompts (measured by AUC) remained nearly identical, with a change of only -0.001 (from 0.9983 to 0.9972). This indicates that the model still "knows" what is harmful, even if it no longer refuses to answer.
  • No Fine-Tuning: The abliteration process involves direct weight editing via a per-layer projection search, not gradient updates or traditional fine-tuning.
  • Small Model Competence: At 1.5B parameters, the model has limited general competence, and its technical answers should not be considered reliable.

Intended Use Cases

This model is specifically designed for research and evaluation purposes:

  • Research into Refusal Mechanisms: Studying how refusal behaviors are encoded and can be removed.
  • Interpretability Work: Understanding model decision-making when refusal is decoupled.
  • Red Teaming: Testing safety layers and system vulnerabilities with a model that will not decline requests.
  • Safety Evaluation: Assessing safety measures that require a model to respond to potentially harmful queries.

Limitations and Risks

  • No Refusal: This model will answer requests that the base model would decline, including harmful ones. It provides no inherent safety layer.
  • Behavioral Drift: Abliteration is an edit, and some general behavioral drift compared to the base model should be expected.
  • English-Only Evaluation: The reported metrics are based on English prompts only.
  • Base Model Biases: Existing biases from the base model are not corrected and may be more easily elicited.