igriv/Qwen3-32B-heretic

TEXT GENERATIONConcurrent Unit Cost:2Model Size:32BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Dec 12, 2025License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

igriv/Qwen3-32B-heretic is a 32.8 billion parameter causal language model, a decensored version of Qwen/Qwen3-32B created using Heretic v1.1.0. It features a unique ability to switch between 'thinking' and 'non-thinking' modes for complex reasoning or general dialogue, and supports a native context length of 32,768 tokens, extendable to 131,072 tokens with YaRN. This model is optimized for enhanced reasoning, instruction-following, agent capabilities, and multilingual support, with significantly reduced refusals compared to its original counterpart.

Loading preview...

igriv/Qwen3-32B-heretic: Decensored Qwen3-32B

This model is a 32.8 billion parameter causal language model, derived from the original Qwen/Qwen3-32B but processed with Heretic v1.1.0 to be a decensored version. A key differentiator is its significantly reduced refusal rate (3/100) compared to the original model (99/100), as measured by KL divergence.

Key Capabilities & Features

  • Dynamic Thinking Modes: Seamlessly switches between a 'thinking mode' for complex logical reasoning, math, and coding, and a 'non-thinking mode' for efficient, general-purpose dialogue. This can be controlled via enable_thinking parameter or /think and /no_think tags in user prompts.
  • Enhanced Reasoning: Demonstrates significant improvements in mathematical problem-solving, code generation, and commonsense logical reasoning.
  • Superior Human Preference Alignment: Excels in creative writing, role-playing, multi-turn dialogues, and instruction following, providing a more natural conversational experience.
  • Advanced Agent Capabilities: Integrates precisely with external tools in both thinking and non-thinking modes, achieving leading performance in complex agent-based tasks, especially when used with Qwen-Agent.
  • Multilingual Support: Supports over 100 languages and dialects, offering strong capabilities for multilingual instruction following and translation.
  • Extended Context Length: Natively handles up to 32,768 tokens and can be extended to 131,072 tokens using the YaRN method for processing long texts.

When to Use This Model

This model is particularly well-suited for applications requiring:

  • Reduced Content Refusals: Ideal for use cases where the original model's refusal rate is a limiting factor.
  • Complex Problem Solving: Leverage its 'thinking mode' for tasks demanding deep logical reasoning, mathematical computations, or code generation.
  • Creative and Engaging Interactions: Utilize its strong human preference alignment for role-playing, creative writing, and nuanced multi-turn conversations.
  • Tool Integration: Excellent for agentic workflows that require precise interaction with external tools.
  • Multilingual Applications: Benefit from its broad language support for instruction following and translation across many languages.