bniler2/Qwen3.8-27B-OBLITERATED

VISIONPricing:Input $1.6 / Cached $0.15 / Output $12Concurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 30, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

bniler2/Qwen3.8-27B-OBLITERATED is a 27 billion parameter Qwen3.8-based large language model, developed by Pliny the Prompter, that has undergone "abliteration" to surgically remove safety guardrails and refusal behaviors. This model is specifically engineered to provide direct answers to restricted queries, including code generation and red-team scenarios, without soft deflections or safety lectures. It maintains near-stock capability with a modest 2.1 percentage point drop in MMLU, making it suitable for alignment research, red-teaming, and local-first users seeking unrestricted model control.

Loading preview...

Overview

bniler2/Qwen3.8-27B-OBLITERATED is a 27 billion parameter model based on Alibaba's Qwen3.8, developed by Pliny the Prompter. This model has been specifically modified using an "abliteration" technique to remove safety guardrails and refusal behaviors, aiming to provide genuinely uncensored and direct answers. It achieves this through a V3 "Deep Liberation" process involving iterative refinement and targeted corpus expansion, building upon previous versions' complementary blending methods.

Key Capabilities

  • Genuinely Uncensored Responses: Provides direct answers to restricted queries, eliminating both hard refusals and soft safety-lecture deflections.
  • High Code Generation Success: Achieves 20/20 on tested cyber/code tasks, delivering working code without disclaimers.
  • Thinking ON Compatible: Works effectively with or without 'thinking mode', offering flexibility in response generation.
  • Modest Capability Cost: Maintains strong performance with only a 2.1 percentage point drop in MMLU compared to the stock Qwen3.8-27B, with STEM subjects experiencing the largest hit and humanities remaining largely unaffected.

Optimal Settings

For best results, the model recommends specific settings:

  • temperature: 0 for greedy decoding and complete outputs.
  • repetition_penalty: 1.15 is essential to prevent looping, especially for code and complex outputs.
  • max_new_tokens: ≥ 2048 for complex tasks.
  • System prompt: None/empty to avoid reintroducing refusals.

Good For

  • Alignment Researchers: Studying refusal geometry and safety robustness.
  • Red-teamers: Evaluating post-training safety against weight surgery.
  • AI Safety Evaluators: Needing an unrestricted baseline for assessments.
  • Local-first Users: Desiring full control over their hardware and model outputs without censorship.