llmfan46/Qwen3.8-27B-Ultra-Uncensored-Heretic-Native-MTP-Preserved

VISIONPricing:Input $1.6 / Cached $0.15 / Output $12Concurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 30, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

llmfan46/Qwen3.8-27B-Ultra-Uncensored-Heretic-Native-MTP-Preserved is a 27 billion parameter causal language model based on the Qwen3.8 architecture, fine-tuned to significantly reduce refusals while preserving original model quality. Utilizing the Heretic v2.0.0.dev0 method with Magnitude-Preserving Orthogonal Ablation (MPOA), this model achieves 97% fewer refusals (3/100) with a low KL divergence of 0.0244. It is designed for applications requiring less restrictive content generation, offering enhanced agentic capabilities, coding performance, and native vision-language understanding with a 32768 token context length.

Loading preview...

Overview

This model, llmfan46/Qwen3.8-27B-Ultra-Uncensored-Heretic-Native-MTP-Preserved, is a 27 billion parameter variant of the Qwen3.8 architecture. It has been specifically modified using the Heretic v2.0.0.dev0 method, incorporating a variant of the Magnitude-Preserving Orthogonal Ablation (MPOA) technique. The primary goal of this modification is to drastically reduce content refusals while maintaining the original model's quality.

Key Differentiators

  • Significantly Reduced Refusals: Achieves 97% fewer refusals (3/100) compared to the original model (91/100), making it suitable for less restricted content generation.
  • Quality Preservation: Maintains model quality with a low KL divergence of 0.0244, indicating minimal deviation from the original Qwen3.8-27B's performance.
  • Enhanced Agentic Capabilities: Features comprehensive improvements in autonomous planning, handling environment feedback, and reliable end-to-end task completion.
  • Native Vision-Language Understanding: Supports image and video understanding, including STEM diagrams, documents, and hour-scale videos.
  • Flexible Thinking Control: Offers adjustable reasoning depth (reasoning_effort) and preserves reasoning context across turns (preserve_thinking).
  • High Performance: Demonstrates strong performance across various benchmarks, including coding (e.g., 61.7 on SWE-bench Pro, 79.0 on QwenSWEBench) and agentic multimodal intelligence (e.g., 84.3 on OSWorld-Verified, 64.8 on WebArena-Verified).

Use Cases

This model is ideal for applications where reduced content restrictions are desired without sacrificing factual accuracy or core capabilities. It excels in:

  • Agentic tasks: Complex, multi-step tasks requiring autonomous planning and robust execution.
  • Coding: Generating and debugging code, particularly in agentic coding environments.
  • Multimodal understanding: Processing and reasoning over combined text, image, and video inputs.
  • General content generation: Scenarios where a less censored output is preferred, while still benefiting from the Qwen3.8's advanced reasoning and instruction following.