richardyoung/Qwen3-4B-Instruct-2507-heretic

TEXT GENERATIONConcurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 24, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The richardyoung/Qwen3-4B-Instruct-2507-heretic is a 4 billion parameter causal language model, based on the Qwen3-4B-Instruct-2507 architecture, that has been decensored using the Heretic v1.4.0 tool. This model retains the original's enhanced capabilities in instruction following, logical reasoning, and long-context understanding up to 262,144 tokens, while significantly reducing refusals compared to the original. It is optimized for open-ended tasks and scenarios requiring less restrictive content generation, making it suitable for diverse applications where the original model's content filters might be too stringent.

Loading preview...

Overview

This model, richardyoung/Qwen3-4B-Instruct-2507-heretic, is a decensored version of the Qwen3-4B-Instruct-2507 model, created using the Heretic v1.4.0 tool. It is a 4 billion parameter causal language model with a native context length of 262,144 tokens. The primary differentiator is its significantly reduced refusal rate (3/100 compared to 100/100 for the original), achieved through the decensoring process, while maintaining the original model's performance.

Key Capabilities

  • Decensored Output: Offers less restrictive content generation compared to its base model.
  • Enhanced General Capabilities: Demonstrates improvements in instruction following, logical reasoning, text comprehension, mathematics, science, coding, and tool usage.
  • Long-Context Understanding: Natively supports a context length of 262,144 tokens.
  • Multilingual Support: Shows substantial gains in long-tail knowledge coverage across multiple languages.
  • Agentic Use: Excels in tool calling capabilities, recommended for use with Qwen-Agent.

Performance Highlights

While maintaining the strong performance of the base Qwen3-4B-Instruct-2507, this 'heretic' version specifically addresses content filtering. The original model showed strong benchmarks across various categories including Knowledge (MMLU-Pro: 69.6), Reasoning (AIME25: 47.4), Coding (MultiPL-E: 76.8), and Alignment (Creative Writing v3: 83.5). The decensoring process aims to provide these capabilities with fewer content restrictions.

Good For

  • Applications requiring less filtered or more open-ended text generation.
  • Scenarios where the original Qwen3-4B-Instruct-2507's content policies are too restrictive.
  • Developers seeking a powerful 4B parameter model with extensive context and strong general AI capabilities, including agentic workflows.