Kewk/Heretical-Qwen3.5-9B
Kewk/Heretical-Qwen3.5-9B is a 9 billion parameter causal language model based on the Qwen3.5 architecture, fine-tuned to significantly reduce refusals. It leverages a hybrid Gated DeltaNet + Softmax Attention design and offers a native context length of 262,144 tokens, extensible up to 1,010,000 tokens. This model excels in multimodal learning, agentic capabilities, and long-context understanding, making it suitable for applications requiring robust, less restrictive AI responses.
Loading preview...
Heretical-Qwen3.5-9B: A Decensored Multimodal Powerhouse
This model, Heretical-Qwen3.5-9B, is a 9 billion parameter variant of the Qwen3.5 architecture, specifically fine-tuned using a custom Heretic fork to drastically reduce refusal rates (3/100 refusals compared to 100/100 in the original model). It maintains the impressive capabilities of its base model, Qwen3.5, which features a hybrid Gated DeltaNet + Softmax Attention architecture.
Key Capabilities & Performance
- Reduced Refusals: Achieves a refusal rate of just 3/100, making it highly suitable for use cases where content filtering is undesirable.
- Multimodal Learning: Offers unified vision-language foundation, outperforming previous Qwen-VL models across reasoning, coding, agents, and visual understanding benchmarks.
- Exceptional Benchmarks: Demonstrates strong performance across various benchmarks, including GPQA Diamond (81.7), MMLU-Pro (82.5), IFEval (91.5), MMMU-Pro (70.1), MathVision (78.9), and VideoMME (84.5), often competitive with or surpassing much larger models.
- Long Context Window: Natively supports a context length of 262,144 tokens, extensible up to 1,010,000 tokens using YaRN scaling techniques.
- Efficient Hybrid Architecture: Utilizes Gated Delta Networks and sparse Mixture-of-Experts for high-throughput inference with minimal latency.
- Global Linguistic Coverage: Expanded support for 201 languages and dialects.
- Agentic Capabilities: Excels in tool calling, with recommendations for use with Qwen-Agent and Qwen Code for building agent applications.
Good For
- Applications requiring less restrictive content generation and responses.
- Complex multimodal tasks involving images and video, such as visual question answering, document understanding, and video summarization.
- Long-context understanding and generation, especially for tasks requiring extensive document processing or multi-turn conversations.
- Agent-based applications and tool use, leveraging its strong instruction following and reasoning abilities.
- Scenarios demanding high performance in knowledge, STEM, and reasoning tasks within a 9B parameter budget.