rAVEUK/Qwen3.8-27B-Uncensored

VISIONPricing:Input $1.6 / Cached $0.15 / Output $12Concurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Sep 14, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

rAVEUK/Qwen3.8-27B-Uncensored is a 27 billion parameter, bf16 precision language model derived from Qwen/Qwen3.8-27B, developed by rAVEUK. It features a 262144 token context length and vision capabilities. This model significantly reduces refusal behavior on harmful prompts from 98% to 12% compared to its base, achieved through the Heretic method without fine-tuning, making it suitable for research and local inference where reduced content moderation is desired.

Loading preview...

Qwen3.8-27B-Uncensored Overview

rAVEUK/Qwen3.8-27B-Uncensored is a 27 billion parameter model based on Qwen/Qwen3.8-27B, designed to substantially reduce refusal behavior while maintaining the original model's core capabilities. It retains the Qwen3_5ForConditionalGeneration architecture, 64 layers, a 248320-token vocabulary, and vision support with a 262144-token context length. The model is provided in bf16 precision.

Key Capabilities & Differentiators

  • Significantly Reduced Refusals: Achieves a reduction from 98/100 refusals on harmful prompts (base model) to 12/100 in this uncensored version. This was accomplished using the Heretic method, which co-minimizes refusal count against KL divergence from the base model, without traditional fine-tuning or additional training data.
  • Minimal Performance Impact: Benchmarks show a mean drop of only -0.5 points across MMLU, ARC-Challenge, HellaSwag, and Winogrande compared to the base model. These deltas are within or close to standard error, indicating that the refusal reduction has a negligible impact on general reasoning abilities.
  • Vision-Capable: Inherits the vision capabilities from the base Qwen3.8-27B model.
  • Multi-Token Prediction (MTP) Head: The MTP head is present and verified, supporting speculative decoding.

Good for

  • Research into Model Behavior: Ideal for studying model responses without aggressive refusal filtering, particularly for exploring the boundaries of content generation.
  • Local Inference: Optimized for local deployment, with GGUF quantizations available for llama.cpp via JonathanColetti/Qwen3.8-27B-Uncensored-GGUF.
  • Applications Requiring Less Censorship: Suitable for use cases where the primary goal is to generate responses with significantly reduced content moderation, provided appropriate safety layers are implemented by the user.