abliterant/Qwen3.8-27B-RANA-abliterated
abliterant/Qwen3.8-27B-RANA-abliterated is a 27.78 billion parameter language model developed by Abliterant, based on Qwen/Qwen3.8-27B, with a configured context length of 262,144 tokens. This model utilizes Reasoning-Anchored, Norm-preserving Ablation (RANA) to remove refusal behaviors, making it a safety-alignment-removed research model. Its primary use case is for research into interpretability, red-teaming, and robustness evaluation of refusal behavior, as it will produce content the base model refuses.
Loading preview...
Qwen3.8-27B-RANA-abliterated: A Refusal-Ablated Research Model
This model, developed by Abliterant, is a 27.78 billion parameter version of Qwen/Qwen3.8-27B that has undergone Reasoning-Anchored, Norm-preserving Ablation (RANA). The RANA method removes a single "refusal direction" from the model's residual stream, effectively eliminating most safety-alignment-driven refusals while preserving the original weight norms. This makes it a safety-alignment-removed research model intended for specific research purposes.
Key Capabilities
- Refusal Behavior Research: Designed to produce content that the base Qwen model would refuse, enabling studies on interpretability, red-teaming, and robustness evaluation of refusal mechanisms.
- Preserved Core Capabilities: Evaluation shows that core capabilities remain largely intact, with an average capability change of 0.82 percentage points versus the base model across various tasks. The largest drop is ~1.5 pp on TruthfulQA.
- Tool Use and Vision: Supports vision input, multi-turn tool calling, and MTP speculative decoding, with these components remaining byte-identical or functionally preserved after ablation.
- Extended Context: Configured for a 262,144-token context, though evaluated up to a 20,480-token serving limit.
Good For
- Academic Research: Ideal for researchers investigating model safety, alignment, and the mechanisms behind refusal behaviors.
- Red-Teaming: Useful for probing model vulnerabilities and generating content that typically triggers safety filters.
- Interpretability Studies: Provides a platform to understand how specific directions in a model's latent space influence its output and alignment.
Important Note: This model is explicitly a safety-alignment-removed research tool and is not intended for public or end-user deployment without external moderation layers. Users are responsible for compliance with applicable laws and the Apache-2.0 license.