jushys/Qwen3.5-4B-Claude-4.6-OS-Auto-Variable-HERETIC-UNCENSORED-THINKING
The jushys/Qwen3.5-4B-Claude-4.6-OS-Auto-Variable-HERETIC-UNCENSORED-THINKING model is a 4.5 billion parameter fine-tuned variant of the Qwen 3.5 4B dense model, developed by jushys. It was trained using four Claude datasets to enhance reasoning and output generation, surpassing the base model's benchmarks. This model is notable for its "HERETIC" and fully uncensored nature, designed to follow user instructions without refusal, and supports vision capabilities with a 32K token context length.
Loading preview...
Model Overview
jushys/Qwen3.5-4B-Claude-4.6-OS-Auto-Variable-HERETIC-UNCENSORED-THINKING is a 4.5 billion parameter model, fine-tuned from the Qwen 3.5 4B dense model. It leverages four Claude datasets to significantly improve reasoning and output generation, outperforming the original Qwen 3.5 4B on various benchmarks. A key differentiator is its "HERETIC" training, meaning it is designed to be fully uncensored and execute user requests without refusals, offering a unique approach to instruction following.
Key Capabilities
- Enhanced Reasoning & Output: Fine-tuning on Claude datasets has improved the model's ability to reason and generate high-quality outputs, exceeding the base model's performance across multiple benchmarks.
- Uncensored & Heretic Thinking: This model is explicitly designed to be uncensored and follow user instructions directly, without built-in safety alignments that might lead to refusals. Its refusal rate is significantly lower (4/100) compared to the original Qwen 3.5 4B (94/100).
- Multimodal Support: The model supports vision (image) input, with testing confirming its functionality. Video portions were not explicitly tested by the developer.
- Tool Handling & Jinja Template: Features an upgraded Jinja template to address common issues like repeats, long thinking loops, and improved tool handling.
- Long Context Window: Supports a native context length of 262,144 tokens, extensible up to 1,010,000 tokens using RoPE scaling techniques like YaRN.
Good For
- Applications requiring uncensored responses: Ideal for use cases where strict adherence to instructions and unfiltered output is paramount.
- Complex reasoning tasks: The Claude dataset fine-tuning makes it suitable for tasks demanding improved logical thinking and comprehensive output.
- Multimodal applications: Its vision capabilities make it a strong candidate for tasks involving image understanding and generation.
- Developers seeking flexible instruction following: The "HERETIC" nature provides a model that prioritizes user intent over predefined safety alignments.