MuXodious/Qwen3.5-4B-SOMPOA-heresy-v2-MTP
MuXodious/Qwen3.5-4B-SOMPOA-heresy-v2-MTP is a 4.5 billion parameter Qwen3.5 fine-tune, developed by MuXodious, featuring Multi Token Prediction (MTP) weights restored from the base model. This model was produced using P-E-W's Heretic engine with Self-Organizing Maps & Magnitude-Preserving Orthogonal Ablation (SOMA and MPOA) for ablation. It offers a 32768 token context length and is notable for its strong performance in general understanding (UGI score) among models under 20B parameters.
Loading preview...
Model Overview
MuXodious/Qwen3.5-4B-SOMPOA-heresy-v2-MTP is a 4.5 billion parameter fine-tuned version of the Qwen3.5-4B model, developed by MuXodious. It incorporates Multi Token Prediction (MTP) capabilities, with weights restored from the original base model. The fine-tuning process utilized P-E-W's Heretic engine (v1.2.0) with Self-Organizing Maps & Magnitude-Preserving Orthogonal Ablation (SOMA and MPOA) as the ablation method.
Key Differentiators
- Ablation Method: Employs a unique ablation technique (SOMA and MPOA) via the Heretic engine, which is noted for its impact on model characteristics, particularly in reducing refusals.
- Multi Token Prediction (MTP): Features restored MTP weights, enhancing its predictive capabilities.
- UGI Score: As of July 2026, this model holds the second-best UGI (User General Intelligence) score among models with 20 billion parameters or fewer, indicating strong general understanding.
- Context Length: Supports a substantial context length of 32768 tokens, extensible up to 1,010,000 tokens using YaRN scaling techniques.
Core Capabilities
- Multimodal Understanding: Inherits Qwen3.5's unified vision-language foundation, excelling in reasoning, coding, agentic tasks, and visual understanding.
- Multilingual Support: Expanded support for 201 languages and dialects, ensuring global accessibility.
- Agentic Functionality: Designed to work effectively with agent frameworks like Qwen-Agent and Qwen Code, supporting tool calling and complex task automation.
Ideal Use Cases
This model is particularly well-suited for applications requiring:
- General-purpose AI tasks where a strong understanding and reduced refusal rate are beneficial.
- Multimodal applications involving both text and visual inputs, including video understanding.
- Agentic workflows and tool-use scenarios, especially within the Qwen-Agent and Qwen Code ecosystems.
- Long-context processing for complex documents or conversations, leveraging its extended context capabilities.