coder3101/Qwen3-VL-8B-Thinking-heretic

VISIONConcurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Nov 23, 2025License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

coder3101/Qwen3-VL-8B-Thinking-heretic is an 8 billion parameter vision-language model, a decensored version of Qwen/Qwen3-VL-8B-Thinking, offering a 32768 token context length. This model provides comprehensive upgrades in text understanding, visual perception, and reasoning, with significantly reduced refusals compared to its original counterpart. It is optimized for advanced multimodal reasoning, visual agent capabilities, and long context video understanding.

Loading preview...

Overview

coder3101/Qwen3-VL-8B-Thinking-heretic is an 8 billion parameter vision-language model, derived from Qwen/Qwen3-VL-8B-Thinking, with a key modification: it has been "decensored" using the Heretic v1.0.1 tool. This process significantly reduces refusals, with the model showing 1 refusal out of 100 compared to 44/100 in the original model, while maintaining a low KL divergence of 0.03.

Key Capabilities

  • Enhanced Multimodal Reasoning: Excels in STEM/Math tasks, providing causal analysis and logical, evidence-based answers.
  • Visual Agent: Capable of operating PC/mobile GUIs by recognizing elements, understanding functions, and completing tasks.
  • Advanced Spatial Perception: Judges object positions, viewpoints, and occlusions, enabling stronger 2D and 3D grounding for spatial reasoning.
  • Long Context & Video Understanding: Features a native 256K context, expandable to 1M, allowing it to handle extensive documents and hours-long video with full recall.
  • Upgraded Visual Recognition & OCR: Recognizes a broad range of entities and supports OCR in 32 languages, robustly handling challenging conditions and complex characters.
  • Visual Coding Boost: Generates Draw.io/HTML/CSS/JS from image and video inputs.

Model Architecture Updates

Key architectural innovations include Interleaved-MRoPE for enhanced long-horizon video reasoning, DeepStack for fusing multi-level ViT features to capture fine-grained details, and Text–Timestamp Alignment for precise event localization in videos.

Good For

This model is particularly well-suited for applications requiring advanced vision-language understanding, multimodal reasoning, and tasks where reduced content moderation or "decensoring" is desired. Its capabilities make it ideal for visual agents, complex spatial analysis, and processing long-form video and text content.