Kewk/Heretical-Qwen3.5-9B

VISIONConcurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Mar 3, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

Kewk/Heretical-Qwen3.5-9B is a 9 billion parameter causal language model based on the Qwen3.5 architecture, fine-tuned to significantly reduce refusals. It leverages a hybrid Gated DeltaNet + Softmax Attention design and offers a native context length of 262,144 tokens, extensible up to 1,010,000 tokens. This model excels in multimodal learning, agentic capabilities, and long-context understanding, making it suitable for applications requiring robust, less restrictive AI responses.

Loading preview...

Heretical-Qwen3.5-9B: A Decensored Multimodal Powerhouse

This model, Heretical-Qwen3.5-9B, is a 9 billion parameter variant of the Qwen3.5 architecture, specifically fine-tuned using a custom Heretic fork to drastically reduce refusal rates (3/100 refusals compared to 100/100 in the original model). It maintains the impressive capabilities of its base model, Qwen3.5, which features a hybrid Gated DeltaNet + Softmax Attention architecture.

Key Capabilities & Performance

  • Reduced Refusals: Achieves a refusal rate of just 3/100, making it highly suitable for use cases where content filtering is undesirable.
  • Multimodal Learning: Offers unified vision-language foundation, outperforming previous Qwen-VL models across reasoning, coding, agents, and visual understanding benchmarks.
  • Exceptional Benchmarks: Demonstrates strong performance across various benchmarks, including GPQA Diamond (81.7), MMLU-Pro (82.5), IFEval (91.5), MMMU-Pro (70.1), MathVision (78.9), and VideoMME (84.5), often competitive with or surpassing much larger models.
  • Long Context Window: Natively supports a context length of 262,144 tokens, extensible up to 1,010,000 tokens using YaRN scaling techniques.
  • Efficient Hybrid Architecture: Utilizes Gated Delta Networks and sparse Mixture-of-Experts for high-throughput inference with minimal latency.
  • Global Linguistic Coverage: Expanded support for 201 languages and dialects.
  • Agentic Capabilities: Excels in tool calling, with recommendations for use with Qwen-Agent and Qwen Code for building agent applications.

Good For

  • Applications requiring less restrictive content generation and responses.
  • Complex multimodal tasks involving images and video, such as visual question answering, document understanding, and video summarization.
  • Long-context understanding and generation, especially for tasks requiring extensive document processing or multi-turn conversations.
  • Agent-based applications and tool use, leveraging its strong instruction following and reasoning abilities.
  • Scenarios demanding high performance in knowledge, STEM, and reasoning tasks within a 9B parameter budget.