p-e-w/Qwen3-4B-Instruct-2507-heretic-v2
p-e-w/Qwen3-4B-Instruct-2507-heretic-v2 is a 4 billion parameter instruction-tuned causal language model, derived from Qwen/Qwen3-4B-Instruct-2507 and modified using Heretic v1.1.0 for decensoring. It features a native context length of 262,144 tokens and demonstrates significant improvements in instruction following, logical reasoning, and long-tail knowledge coverage. This model is particularly optimized for subjective and open-ended tasks, offering enhanced capabilities in areas like creative writing and agentic tool usage.
Loading preview...
Overview
This model, p-e-w/Qwen3-4B-Instruct-2507-heretic-v2, is a 4 billion parameter instruction-tuned causal language model. It is a decensored version of the original Qwen/Qwen3-4B-Instruct-2507, created using the Heretic v1.1.0 tool. The model maintains a substantial native context length of 262,144 tokens, though a recommended operational context length of 16,384 tokens is suggested for most queries to avoid out-of-memory issues.
Key Capabilities & Enhancements
- Decensored Output: Modified to reduce refusals, showing 8/100 refusals compared to 100/100 in the original model.
- General Capabilities: Significant improvements across instruction following, logical reasoning, text comprehension, mathematics, science, coding, and tool usage.
- Knowledge Coverage: Substantial gains in long-tail knowledge coverage, including across multiple languages.
- User Alignment: Markedly better alignment with user preferences for subjective and open-ended tasks, leading to more helpful responses and higher-quality text generation.
- Long-Context Understanding: Enhanced capabilities in processing and understanding long contexts up to 256K tokens.
- Agentic Use: Excels in tool calling, with recommendations to use Qwen-Agent for optimal performance.
Performance Highlights
Compared to its base model and other benchmarks, Qwen3-4B-Instruct-2507-heretic-v2 demonstrates strong performance across various metrics:
- Knowledge: Achieves 69.6 on MMLU-Pro and 84.2 on MMLU-Redux.
- Reasoning: Scores 47.4 on AIME25 and 80.2 on ZebraLogic.
- Coding: Reaches 35.1 on LiveCodeBench v6 and 76.8 on MultiPL-E.
- Alignment: Shows 83.4 on IFEval and 83.5 on Creative Writing v3.
- Agent: Achieves 61.9 on BFCL-v3 and 48.7 on TAU1-Retail.
Recommended Use Cases
This model is well-suited for applications requiring:
- Open-ended text generation and creative writing.
- Complex instruction following and logical reasoning tasks.
- Tool-use and agentic workflows due to its strong tool-calling capabilities.
- Processing and understanding very long documents or conversations.
- Scenarios where a less restrictive or decensored output is desired.