saidutta69/Qwen3-4B-Instruct-2507-heretic
The saidutta69/Qwen3-4B-Instruct-2507-heretic is a 4.02 billion parameter, instruction-tuned causal language model based on the Qwen3 architecture, developed by saidutta69. This variant is a decensored version of Qwen/Qwen3-4B-Instruct-2507, created using Heretic v1.4.0 for targeted refusal suppression. It features a 256K context length and is optimized for instruction following, tool use, and long-context tasks, making it suitable for latency-sensitive agent loops and structured data extraction.
Loading preview...
Qwen3-4B-Instruct-2507-heretic: Decensored Instruction Following
This model is a 4.02 billion parameter, instruction-tuned variant of the Qwen3-4B-Instruct-2507 base model, developed by saidutta69. It has been decensored using the Heretic v1.4.0 "abliteration" technique, which involves targeted weight edits to suppress refusal behavior without fine-tuning. This approach maintains the original instruction-following and tool-use capabilities of the 2507 refresh while significantly reducing refusals.
Key Capabilities
- Refusal Suppression: Refusals drop from 100/100 to 3/100, effectively removing most refusal behaviors present in the base model.
- Instruction Following: Retains the strong instruction-following, logical reasoning, and text comprehension improvements from the Qwen3-4B-Instruct-2507 refresh.
- Tool Use & Long Context: Optimized for tool calling, structured extraction, and agent scaffolding, benefiting from its 256K context length and improved tool-usage accuracy.
- Efficiency: At 4B parameters, it is designed for fast token generation, making it suitable for latency-sensitive applications.
- Accessibility: Provided with a full suite of GGUF quantizations, enabling efficient deployment on CPU and low-VRAM GPUs.
Good For
- Developers requiring a fast 4B model that strictly follows instructions without generating refusals.
- Applications involving tool-calling loops, structured data extraction, and agent scaffolding where per-token latency is critical.
- Use cases demanding long-context understanding (up to 256K tokens) for complex tasks.
- Environments where the base model's refusal behavior is undesirable, and a more compliant model is needed.
Note: This model's refusal suppression is deliberate. It will comply with requests the base model would refuse, including some that may be considered unsafe. Users are responsible for its deployment and should not use it behind unmoderated public-facing endpoints.