MuXodious/Qwen3.5-4B-PaperWitch-heresy-v2

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Mar 2, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

MuXodious/Qwen3.5-4B-PaperWitch-heresy-v2 is a 4.5 billion parameter Qwen3.5-4B fine-tune developed by MuXodious, optimized using P-E-W's Heretic ablation engine with Magnitude-Preserving Orthogonal Ablation. This model specifically targets reducing disclaimers and improving one-shot prompting, while maintaining a 32768 token context length. It is designed to enhance directness in responses, making it suitable for applications requiring less cautious or verbose outputs.

Loading preview...

Model Overview

MuXodious/Qwen3.5-4B-PaperWitch-heresy-v2 is a 4.5 billion parameter fine-tuned variant of the Qwen3.5-4B model, developed by MuXodious. This version was created using P-E-W's Heretic (v1.2.0) ablation engine, specifically with Magnitude-Preserving Orthogonal Ablation, to modify its response characteristics.

Key Differentiators

  • Reduced Disclaimers: The primary goal of this fine-tune is to lessen the model's tendency to add disclaimers, aiming for more direct and concise outputs.
  • Improved One-Shot Prompting: By reducing secondary factors that lead to verbose or cautious responses, the model is intended to perform better in one-shot prompting scenarios.
  • Ablation-Based Tuning: Utilizes a unique "heretication" process to achieve specific behavioral changes, focusing on refusal rates and KL divergence.
  • Multimodal Capabilities: Inherits the Qwen3.5 base model's unified vision-language foundation, supporting image and video inputs, and features an efficient hybrid architecture with Gated Delta Networks and sparse Mixture-of-Experts.
  • Extended Context Length: Natively supports a context length of 262,144 tokens, extensible up to 1,010,000 tokens using YaRN scaling techniques.

Performance Insights

While the ablation process virtually increases refusals and KL divergence compared to the base Qwen3.5, it successfully reduces the model's disclaimer tendency. The underlying Qwen3.5-4B demonstrates strong performance across various benchmarks, including:

  • Knowledge & STEM: Achieves 79.1 on MMLU-Pro and 85.1 on C-Eval.
  • Instruction Following: Scores 89.8 on IFEval.
  • Vision Language: Attains 77.6 on MMMU and 89.4 on MMBench.
  • Multilingualism: Supports 201 languages and dialects, with 71.5 on MMLU-ProX.

Ideal Use Cases

  • Applications requiring direct responses: Suitable for scenarios where a model's tendency to add disclaimers is undesirable.
  • One-shot prompting: Optimized for tasks that benefit from concise, immediate answers without extensive caveats.
  • Multimodal tasks: Leverages the base Qwen3.5's capabilities for processing and understanding image and video inputs.
  • Agentic applications: Can be integrated with frameworks like Qwen-Agent and Qwen Code for tool-calling and code-related tasks.