richardyoung/Qwen3-4B-Instruct-2507-heretic

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 24, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

The richardyoung/Qwen3-4B-Instruct-2507-heretic is a 4 billion parameter instruction-tuned causal language model, based on the Qwen3-4B-Instruct-2507 architecture, developed by Richard Young. This model is a decensored version, created using the Heretic v1.4.0 tool, specifically designed to reduce refusals compared to the original Qwen model. It features a native context length of 262,144 tokens and is optimized for general capabilities including instruction following, logical reasoning, and coding, with a focus on providing more helpful and less restrictive responses.

Loading preview...

Model Overview

This model, richardyoung/Qwen3-4B-Instruct-2507-heretic, is a 4 billion parameter instruction-tuned causal language model derived from the Qwen3-4B-Instruct-2507 base model. Developed by Richard Young, its primary distinction is being a decensored version, achieved using the Heretic v1.4.0 tool. This modification significantly reduces the model's refusal rate, offering a more open and less restrictive response generation compared to the original Qwen model.

Key Capabilities & Enhancements

  • Decensored Output: Achieves a refusal rate of 3/100 compared to the original model's 100/100, making it suitable for use cases requiring fewer content restrictions.
  • General Capabilities: Inherits and enhances the Qwen3-4B-Instruct-2507's improvements in instruction following, logical reasoning, text comprehension, mathematics, science, coding, and tool usage.
  • Long Context Understanding: Supports a native context length of 262,144 tokens, enabling processing of extensive inputs.
  • Subjective & Open-Ended Tasks: Demonstrates improved alignment with user preferences for more helpful and higher-quality text generation in subjective and open-ended scenarios.
  • Tool Calling: Excels in tool calling capabilities, recommended for use with Qwen-Agent for agentic applications.

Performance Highlights

Compared to the original Qwen3-4B-Instruct-2507, this model maintains strong performance across various benchmarks, while specifically addressing content moderation. For instance, it shows competitive scores in MMLU-Pro (69.6), AIME25 (47.4), and LiveCodeBench v6 (35.1), often outperforming the base Qwen3-4B Non-Thinking model.

When to Use This Model

  • Use Cases: Ideal for applications where a less restrictive and more direct response generation is preferred, particularly in creative writing, open-ended dialogue, or scenarios where the original model's content filters might be too aggressive.
  • Considerations: While decensored, users should be mindful of the content generated and implement appropriate safeguards for their specific applications. The model supports non-thinking mode exclusively, simplifying its use without requiring enable_thinking=False.