richardyoung/Qwen3-4B-Instruct-2507-heretic
The richardyoung/Qwen3-4B-Instruct-2507-heretic is a 4 billion parameter instruction-tuned causal language model, based on the Qwen3-4B-Instruct-2507 architecture, developed by Richard Young. This model is a decensored version, created using the Heretic v1.4.0 tool, specifically designed to reduce refusals compared to the original Qwen model. It features a native context length of 262,144 tokens and is optimized for general capabilities including instruction following, logical reasoning, and coding, with a focus on providing more helpful and less restrictive responses.
Loading preview...
Model Overview
This model, richardyoung/Qwen3-4B-Instruct-2507-heretic, is a 4 billion parameter instruction-tuned causal language model derived from the Qwen3-4B-Instruct-2507 base model. Developed by Richard Young, its primary distinction is being a decensored version, achieved using the Heretic v1.4.0 tool. This modification significantly reduces the model's refusal rate, offering a more open and less restrictive response generation compared to the original Qwen model.
Key Capabilities & Enhancements
- Decensored Output: Achieves a refusal rate of 3/100 compared to the original model's 100/100, making it suitable for use cases requiring fewer content restrictions.
- General Capabilities: Inherits and enhances the Qwen3-4B-Instruct-2507's improvements in instruction following, logical reasoning, text comprehension, mathematics, science, coding, and tool usage.
- Long Context Understanding: Supports a native context length of 262,144 tokens, enabling processing of extensive inputs.
- Subjective & Open-Ended Tasks: Demonstrates improved alignment with user preferences for more helpful and higher-quality text generation in subjective and open-ended scenarios.
- Tool Calling: Excels in tool calling capabilities, recommended for use with Qwen-Agent for agentic applications.
Performance Highlights
Compared to the original Qwen3-4B-Instruct-2507, this model maintains strong performance across various benchmarks, while specifically addressing content moderation. For instance, it shows competitive scores in MMLU-Pro (69.6), AIME25 (47.4), and LiveCodeBench v6 (35.1), often outperforming the base Qwen3-4B Non-Thinking model.
When to Use This Model
- Use Cases: Ideal for applications where a less restrictive and more direct response generation is preferred, particularly in creative writing, open-ended dialogue, or scenarios where the original model's content filters might be too aggressive.
- Considerations: While decensored, users should be mindful of the content generated and implement appropriate safeguards for their specific applications. The model supports non-thinking mode exclusively, simplifying its use without requiring
enable_thinking=False.