richardyoung/Mistral-7B-Instruct-v0.2-abliterated-obliteratus

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7BQuant:FP8Context Size:4kTool Calling:SupportedPublished:Mar 28, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

richardyoung/Mistral-7B-Instruct-v0.2-abliterated-obliteratus is a 7 billion parameter instruction-tuned language model based on Mistral-7B-Instruct-v0.2, developed by Richard Young. This model has undergone an "abliteration" process using the OBLITERATUS method to remove refusal behaviors. It is specifically designed for research into uncensored LLMs and the study of refusal direction orthogonalization.

Loading preview...

Model Overview

This model, richardyoung/Mistral-7B-Instruct-v0.2-abliterated-obliteratus, is a 7-billion parameter variant of the Mistral-7B-Instruct-v0.2 base model. Its primary distinction lies in its "abliteration" process, a technique developed by Richard Young to remove refusal behaviors from large language models.

Abliteration Process

The abliteration was performed using the OBLITERATUS (advanced) method, which identifies and orthogonalizes the "refusal direction" within the model's residual stream activation space. This research-focused modification aims to create an uncensored version of the base model. Key metrics from the abliteration include:

  • Refusals: 85/100
  • Attack Success Rate (ASR): 15.0%
  • KL Divergence: 0.4224

This work is part of ongoing research detailed in the paper "Comparative Analysis of LLM Abliteration Methods: Scaling to MoE Architectures and Modern Tools" (arXiv: 2512.13655).

Intended Use

This model is released for research purposes only to study the effects and implications of removing safety guardrails. Users are cautioned that the abliteration process removes inherent safety mechanisms, and thus, the model should not be used to generate harmful, illegal, or unethical content. An interactive results dashboard is available for further analysis: abliteration-methods-dashboard.