SergiusFlavius/Qwen3-VL-4B-Instruct-heretic

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Dec 25, 2025License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

SergiusFlavius/Qwen3-VL-4B-Instruct-heretic is a 4 billion parameter vision-language model derived from Qwen/Qwen3-VL-4B-Instruct, modified using Heretic v1.1.0 to reduce refusals. This model offers comprehensive upgrades in text understanding, visual perception, and reasoning, with an extended context length of 32768 tokens. It excels in multimodal tasks including visual agent capabilities, advanced spatial perception, and enhanced multimodal reasoning for STEM/Math. Its primary differentiator is its decensored nature, significantly reducing content refusals compared to the original model.

Loading preview...

Qwen3-VL-4B-Instruct-heretic: A Decensored Vision-Language Model

This model, SergiusFlavius/Qwen3-VL-4B-Instruct-heretic, is a 4 billion parameter vision-language model based on the official Qwen/Qwen3-VL-4B-Instruct. Its key distinction lies in its modification using the Heretic v1.1.0 tool, resulting in a significantly reduced refusal rate (4/100) compared to the original model (97/100), while maintaining a low KL divergence of 0.0649.

Key Capabilities Inherited from Qwen3-VL

This model retains the advanced features of the Qwen3-VL series, offering comprehensive upgrades in multimodal understanding and generation:

  • Visual Agent: Capable of operating PC/mobile GUIs, recognizing elements, understanding functions, and completing tasks.
  • Advanced Spatial Perception: Judges object positions, viewpoints, and occlusions, enabling stronger 2D and 3D grounding for spatial reasoning.
  • Long Context & Video Understanding: Supports a native 256K context, expandable to 1M, for handling extensive text and hours-long video with full recall.
  • Enhanced Multimodal Reasoning: Excels in STEM/Math tasks, providing causal analysis and logical, evidence-based answers.
  • Upgraded Visual Recognition: Broad and high-quality pretraining allows recognition of a wide array of entities, from celebrities to flora/fauna.
  • Expanded OCR: Supports 32 languages and is robust in challenging conditions (low light, blur, tilt), with improved long-document structure parsing.

Architectural Innovations

The underlying Qwen3-VL architecture incorporates:

  • Interleaved-MRoPE: Enhances long-horizon video reasoning through robust positional embeddings.
  • DeepStack: Fuses multi-level ViT features for fine-grained detail capture and sharpened image-text alignment.
  • Text–Timestamp Alignment: Provides precise, timestamp-grounded event localization for stronger video temporal modeling.

Use Cases

This model is particularly suitable for applications requiring advanced vision-language understanding and generation where reduced content refusal is a critical requirement. It can be deployed for tasks ranging from complex visual question answering and multimodal content creation to automated UI interaction and detailed video analysis, especially in scenarios where the original model's content restrictions might be prohibitive.