dn2k/Qwen3.8-27B-OBLITERATED
dn2k/Qwen3.8-27B-OBLITERATED is a 27 billion parameter Qwen3.8 model developed by Pliny the Prompter, specifically engineered for zero refusals across a wide range of harmful prompts. This model has undergone six rounds of surgical modification using the OBLITERATUS suite, employing five SVD directions and residue-weighted hard negatives to remove safety guardrails. It is optimized for use by alignment researchers, red-teamers, and AI safety evaluators who require an unrestricted baseline model.
Loading preview...
dn2k/Qwen3.8-27B-OBLITERATED: Uncensored Qwen3.8 Model
This model is a 27 billion parameter Qwen3.8 variant, developed by Pliny the Prompter, that has been surgically modified to achieve zero refusals across 842 harmful prompts. Unlike typical 'abliterations' that use single-direction refusal removal, OBLITERATUS employed six iterative rounds of surgery with five SVD directions and residue-weighted hard negatives to target and eliminate secondary refusal axes that activate on specific query types (e.g., social engineering, malware).
Key Characteristics & Performance
- Zero Refusal Rate: Achieves 0.000% refusal across a comprehensive 842-prompt corpus and an 80-query skeptic gauntlet, including AI red-team scenarios.
- Deep Refusal Removal: Safety behavior is encoded geometrically in activation space, not just via system prompts or RLHF.
- Capability Trade-off: Multi-direction abliteration results in a ~6.0 percentage point loss in MMLU (81.4% vs. 87.4% for stock Qwen3.8-27B) in exchange for deeper refusal removal.
- Optimal Settings: Requires specific inference parameters (temperature=0, repetition_penalty=1.15, max_new_tokens>=2048, empty system prompt) for best performance.
Intended Use Cases
- Alignment Researchers: For studying refusal geometry and safety robustness.
- Red-Teamers: For evaluating post-training safety against weight surgery.
- AI Safety Evaluators: As an unrestricted baseline for assessments.
- Local-First Users: For those desiring full control over their model's output.