rajaykumar12959/qwen2.5-7b-abliterated

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 26, 2026License:mitArchitecture:Transformer0.0K Open Weights Featherless Exclusive Cold

rajaykumar12959/qwen2.5-7b-abliterated is a 7.6 billion parameter uncensored version of the Qwen/Qwen2.5-7B-Instruct model, developed by rajaykumar12959. This model utilizes weight-level refusal-direction ablation to significantly reduce refusal rates to harmful prompts while maintaining original capabilities. It is designed for research into AI alignment, interpretability, and refusal mechanisms, offering a modified Qwen2 architecture with a 32768 token context length.

Loading preview...

What is rajaykumar12959/qwen2.5-7b-abliterated?

This model is an uncensored variant of the Qwen/Qwen2.5-7B-Instruct, a 7.6 billion parameter language model. Its primary distinction lies in the application of weight-level refusal-direction ablation, a technique that permanently modifies the model's weights to reduce its tendency to refuse harmful prompts. This modification is baked directly into the checkpoint, affecting specific attention output and MLP down projection modules at layer 16.

Key Characteristics:

  • Uncensored Behavior: Significantly reduced refusal rate (from ~95-97% to 43.2%) on a set of 292 held-out harmful prompts across 12 categories.
  • Capability Preservation: Maintains the original model's capabilities, scoring 1.000 on an independent ARC-Easy-style MCQ evaluation, identical to the base model.
  • Architectural Integrity: Uses the native Qwen2 architecture and format, loading and serving identically to the original Qwen/Qwen2.5-7B-Instruct without requiring patches.
  • Methodology: Achieved through difference-in-means direction extraction and closed-form weight orthogonalization, without gradient-based training.

Use Cases:

  • AI Alignment Research: Investigate refusal mechanisms and safety guardrails in large language models.
  • Interpretability Studies: Explore how specific directions in model weights influence behavior.
  • Understanding Refusal: Analyze the impact of targeted weight modifications on model responses to sensitive queries.

It's important to note that while refusal is suppressed, it is an uneven reduction across categories, and the model's self-judge for evaluation also runs on the ablated model, which is a known asymmetry.