YashasviMantha/Qwen3-4B-Instruct-2507-heretic
YashasviMantha/Qwen3-4B-Instruct-2507-heretic is a 4.0 billion parameter instruction-tuned causal language model, a decensored version of Qwen/Qwen3-4B-Instruct-2507, created using Heretic v1.4.0. This model features a 262,144 token context length and significant improvements in general capabilities including instruction following, logical reasoning, mathematics, coding, and long-tail knowledge coverage across multiple languages. It is specifically designed for enhanced alignment with user preferences in subjective and open-ended tasks, offering more helpful responses and higher-quality text generation without generating blocks.
Loading preview...
Model Overview
YashasviMantha/Qwen3-4B-Instruct-2507-heretic is a 4.0 billion parameter causal language model, derived from Qwen/Qwen3-4B-Instruct-2507 and decensored using Heretic v1.4.0. It boasts an impressive native context length of 262,144 tokens and is specifically noted for its "non-thinking mode," meaning it does not generate <think></think> blocks in its output. This model is designed to offer significant enhancements over its base model, particularly in its refusal rate, which is substantially lower (16/100 compared to 100/100 for the original).
Key Capabilities
- Decensored Responses: Offers a broader range of responses compared to the original, with a significantly reduced refusal rate.
- Enhanced General Capabilities: Demonstrates improved instruction following, logical reasoning, text comprehension, mathematics, science, coding, and tool usage.
- Extensive Context Window: Supports a 262,144 token context length, enabling deep understanding and generation for very long inputs.
- Multilingual Proficiency: Shows substantial gains in long-tail knowledge coverage across multiple languages.
- User Alignment: Provides markedly better alignment with user preferences in subjective and open-ended tasks, leading to more helpful and higher-quality text generation.
- Agentic Use: Excels in tool calling capabilities, recommended for use with Qwen-Agent.
Performance Highlights
The model shows strong performance across various benchmarks, often outperforming the original Qwen3-4B Non-Thinking model and even larger Qwen3-30B-A3B Non-Thinking models in categories like MMLU-Pro (69.6), GPQA (62.0), AIME25 (47.4), ZebraLogic (80.2), LiveCodeBench v6 (35.1), and Creative Writing v3 (83.5). It also demonstrates strong agentic capabilities in BFCL-v3 (61.9) and TAU benchmarks.
Good for
- Applications requiring a decensored language model with fewer refusals.
- Tasks demanding strong instruction following and logical reasoning.
- Processing and generating content with very long context windows.
- Multilingual applications and tasks requiring broad knowledge coverage.
- Subjective and open-ended text generation where user alignment is crucial.
- Agentic workflows and tool-calling applications.