JSon-AI/Qwen3-VL-4B-Instruct-Abliterated-heretic
The JSon-AI/Qwen3-VL-4B-Instruct-Abliterated-heretic is a 4 billion parameter vision-language model, derived from Qwen/Qwen3-VL-4B-Instruct, that has been decensored using the Heretic v1.0.1 tool. This model features enhanced visual perception, reasoning, and extended context length, making it suitable for applications requiring less restrictive content generation. It demonstrates significantly reduced refusals compared to its original counterpart, while maintaining strong multimodal and text understanding capabilities.
Loading preview...
Model Overview
This model, JSon-AI/Qwen3-VL-4B-Instruct-Abliterated-heretic, is a 4 billion parameter vision-language model based on the Qwen3-VL-4B-Instruct architecture. It has been specifically modified using the Heretic v1.0.1 tool to be a decensored version, aiming to reduce content refusal rates.
Key Differentiators
- Decensored Output: Achieves a refusal rate of 6/100 compared to the original model's 92/100, making it suitable for use cases requiring less content filtering.
- Advanced Multimodal Capabilities: Inherits the comprehensive upgrades of the Qwen3-VL series, including superior text understanding and generation, deep visual perception and reasoning, and enhanced spatial and video dynamics comprehension.
- Extended Context Length: Supports a native 256K context, expandable to 1M, allowing it to process extensive textual and visual information, including hours-long video with full recall.
- Visual Agent & Coding Boost: Capable of operating PC/mobile GUIs, recognizing elements, and generating code (Draw.io/HTML/CSS/JS) from images/videos.
- Enhanced Multimodal Reasoning: Excels in STEM/Math tasks, providing causal analysis and logical, evidence-based answers.
Good For
- Applications where reduced content moderation or 'censorship' is desired.
- Tasks requiring advanced visual understanding, such as image description, object recognition, and spatial reasoning.
- Developing visual agents for GUI interaction and automation.
- Code generation from visual inputs.
- Processing and reasoning over long-form video content and extensive documents.