RBergBauer/Qwen3.5-4B-MTP-Heretic

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 25, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

RBergBauer/Qwen3.5-4B-MTP-Heretic is a 4.5 billion parameter Qwen3.5-4B model with a context length of 32768 tokens, specifically modified to ablate its refusal behavior while preserving its multi-token-prediction (MTP) head. This variant significantly reduces safety filtering, making it suitable for research into model alignment and red-teaming. It maintains its original vision-language capabilities, distinguishing it from other ablated Qwen3.5-4B versions that silently drop the MTP head.

Loading preview...

Overview

This model, RBergBauer/Qwen3.5-4B-MTP-Heretic, is a modified version of the Qwen/Qwen3.5-4B base model. Its primary distinction is the ablation of refusal behavior using the heretic tool (v1.4.0), significantly reducing its safety filtering. Crucially, unlike other ablated Qwen3.5-4B variants, this model preserves its multi-token-prediction (MTP) head, ensuring that speculative decoding remains functional and its full vision-language capabilities are retained.

Key Modifications & Features

  • Refusal Ablation: The model's refusal behavior has been reduced from 99/100 to 13/100 on the mlabonne/harmful_behaviors dataset, achieved by orthogonalizing attn.o_proj and mlp.down_proj output projections against a low-rank "refusal direction." This is a weight orthogonalization, not retraining.
  • MTP Head Preservation: The 15 tensors comprising the MTP head, often silently dropped in other heretic exports, have been explicitly re-injected. This ensures the model remains a complete vision-language model with an intact drafter for speculative decoding.
  • Transparency: The exact heretic configuration, trial selection (trial 109), and measured refusal rates are documented, providing a verifiable ablation process.
  • Architecture: It retains the Qwen3_5ForConditionalGeneration architecture, including its vision capabilities, with only the o_proj / down_proj weights of the language model differing from the base.
  • Mixed Precision: The model uses mixed precision, with 723 tensors in float16 and the 15 re-injected MTP tensors in bfloat16, handled transparently by from_pretrained.

Intended Use Cases

  • Research: Ideal for research into refusal directions, model alignment, and understanding safety mechanisms.
  • Experimentation: Suitable for local experimentation with reduced safety alignment.
  • Red-teaming: Useful for red-teaming and developing evaluation harnesses due to its reduced safety filters.

Note: This model has significantly reduced safety filtering and will comply with requests that a stock Qwen3.5-4B refuses, including potentially harmful ones. It is not intended for deployment where untrusted users can interact with it without an additional filtering layer.