jbeslt/Qwen3-8B-heretic

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Nov 28, 2025License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

jbeslt/Qwen3-8B-heretic is an 8.2 billion parameter causal language model, a decensored version of Qwen/Qwen3-8B. It is fine-tuned using the Heretic v1.0.1 method, significantly reducing refusals from 83/100 to 7/100 compared to the original model. This model maintains Qwen3's advanced reasoning, instruction-following, and agent capabilities, with a native context length of 32,768 tokens, extendable to 131,072 tokens with YaRN.

Loading preview...

Overview

jbeslt/Qwen3-8B-heretic is an 8.2 billion parameter large language model, derived from the Qwen/Qwen3-8B base model. This version has been specifically modified using the Heretic v1.0.1 method to significantly reduce content refusals, making it a 'decensored' variant. While the original Qwen3-8B exhibited 83 refusals out of 100, this Heretic version demonstrates only 7 refusals, with a KL divergence of 0.06 from the original.

Key Capabilities

  • Decensored Output: Substantially reduced content refusals compared to the base model, offering more permissive generation.
  • Dual Thinking Modes: Inherits Qwen3's unique ability to seamlessly switch between a 'thinking mode' for complex logical reasoning, math, and coding, and a 'non-thinking mode' for efficient general dialogue. This can be controlled via enable_thinking parameter or /think and /no_think tags in prompts.
  • Enhanced Reasoning: Delivers strong performance in mathematics, code generation, and commonsense logical reasoning, particularly in its thinking mode.
  • Agentic Functions: Excels in tool calling and integration with external tools, achieving leading performance in complex agent-based tasks among open-source models.
  • Multilingual Support: Supports over 100 languages and dialects with robust multilingual instruction following and translation capabilities.
  • Extended Context Window: Natively handles a context length of 32,768 tokens, which can be extended up to 131,072 tokens using the YaRN method.

When to Use This Model

This model is ideal for applications requiring a powerful 8B parameter model with advanced reasoning and agent capabilities, but with a significantly reduced tendency for content refusal. It is particularly suited for use cases where the base Qwen3-8B's refusal rate might be restrictive, while still benefiting from its core strengths in complex problem-solving, code generation, and multilingual interactions. Developers should consider this model when seeking a more unconstrained output from a Qwen3 architecture.