trohrbaugh/Qwen3.6-27B-heretic-ara

VISIONConcurrent Unit Cost:2Model Size:27BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Apr 26, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

trohrbaugh/Qwen3.6-27B-heretic-ara is a 27 billion parameter causal language model, a decensored version of Qwen/Qwen3.6-27B created using the Heretic tool with Arbitrary-Rank Ablation (ARA) method. This model significantly reduces refusals compared to its base, from 99/100 to 4/100, while maintaining strong performance in agentic coding, knowledge, and STEM & reasoning tasks. It supports a native context length of 262,144 tokens, extensible up to 1,010,000 tokens, and is vision-capable, making it suitable for complex multimodal and coding agent applications requiring less restrictive content generation.

Loading preview...

Model Overview

trohrbaugh/Qwen3.6-27B-heretic-ara is a 27 billion parameter causal language model, derived from Qwen/Qwen3.6-27B. This version has been decensored using the Heretic tool with the Arbitrary-Rank Ablation (ARA) method, drastically reducing content refusals from 99/100 in the original model to 4/100, as measured by internal metrics. It maintains the core capabilities of the Qwen3.6 series, including a native context length of 262,144 tokens, extensible up to 1,010,000 tokens, and a vision encoder for multimodal understanding.

Key Capabilities

  • Decensored Output: Significantly reduced refusal rates for broader content generation.
  • Agentic Coding: Enhanced handling of frontend workflows and repository-level reasoning, with strong performance on benchmarks like SWE-bench Verified (77.2%) and Terminal-Bench 2.0 (59.3%).
  • Thinking Preservation: Supports retaining reasoning context from historical messages, improving iterative development and decision consistency.
  • Multimodal Understanding: Processes image and video inputs, excelling in tasks like MMMU (82.9%) and RealWorldQA (84.1%).
  • Ultra-Long Context: Natively supports 262,144 tokens, extendable to over 1 million tokens using RoPE scaling techniques like YaRN.

Good For

  • Applications requiring a less restrictive content policy.
  • Complex coding agent scenarios, including frontend development and repository-level analysis.
  • Multimodal applications involving image and video understanding.
  • Tasks benefiting from extended context windows and preserved reasoning traces.