hpnyaggerman/Qwen3.5-4B-heretic
hpnyaggerman/Qwen3.5-4B-heretic is a 4.5 billion parameter decensored version of the Qwen3.5-4B multimodal language model, developed by Qwen and modified using Heretic v2.0.0.dev0. This model features a 32768 token context length and is optimized for reduced refusals compared to its original counterpart, making it suitable for applications requiring less restrictive content generation. It integrates multimodal learning, architectural efficiency, and scalable reinforcement learning for diverse tasks including vision-language understanding and agentic usage.
Loading preview...
What is hpnyaggerman/Qwen3.5-4B-heretic?
This model is a decensored version of the Qwen3.5-4B multimodal language model, created using Heretic v2.0.0.dev0. It maintains the core capabilities of the original Qwen3.5-4B, which include a 4.5 billion parameter architecture and a native context length of 32,768 tokens, extensible up to 1,010,000 tokens using YaRN scaling.
Key Differentiators
- Decensored Output: Significantly reduces refusals (60/100 compared to 99/100 in the original model), offering less restricted content generation.
- Multimodal Capabilities: Supports unified vision-language understanding, processing both text and image inputs, and video understanding.
- Efficient Hybrid Architecture: Utilizes Gated Delta Networks and sparse Mixture-of-Experts for high-throughput inference with minimal latency.
- Extensive Language Support: Expanded to 201 languages and dialects, ensuring broad global applicability.
- Agentic Features: Excels in tool calling, with recommended integration via Qwen-Agent and Qwen Code for terminal-based AI agent applications.
Should I use this for my use case?
Good for:
- Applications requiring a multimodal model with fewer content restrictions.
- Tasks involving complex reasoning, coding, and agentic behaviors where the original Qwen3.5-4B might refuse to respond.
- Multilingual applications needing broad language support.
- Use cases benefiting from long context processing, especially with the YaRN scaling for ultra-long texts.
Considerations:
- While decensored, users should be aware of potential biases or inappropriate content generation that might arise from reduced safety filters.
- Performance on specific benchmarks for the 'heretic' version is primarily indicated by reduced refusals, while other metrics align with the base Qwen3.5-4B model.