RBergBauer/Qwen3.5-4B-MTP-Heretic
RBergBauer/Qwen3.5-4B-MTP-Heretic is a 4.5 billion parameter Qwen3.5-4B model with a context length of 32768 tokens, specifically modified to ablate its refusal behavior while preserving its multi-token-prediction (MTP) head. This variant significantly reduces safety filtering, making it suitable for research into model alignment and red-teaming. It maintains its original vision-language capabilities, distinguishing it from other ablated Qwen3.5-4B versions that silently drop the MTP head.
Loading preview...
Overview
This model, RBergBauer/Qwen3.5-4B-MTP-Heretic, is a modified version of the Qwen/Qwen3.5-4B base model. Its primary distinction is the ablation of refusal behavior using the heretic tool (v1.4.0), significantly reducing its safety filtering. Crucially, unlike other ablated Qwen3.5-4B variants, this model preserves its multi-token-prediction (MTP) head, ensuring that speculative decoding remains functional and its full vision-language capabilities are retained.
Key Modifications & Features
- Refusal Ablation: The model's refusal behavior has been reduced from 99/100 to 13/100 on the
mlabonne/harmful_behaviorsdataset, achieved by orthogonalizingattn.o_projandmlp.down_projoutput projections against a low-rank "refusal direction." This is a weight orthogonalization, not retraining. - MTP Head Preservation: The 15 tensors comprising the MTP head, often silently dropped in other
hereticexports, have been explicitly re-injected. This ensures the model remains a complete vision-language model with an intact drafter for speculative decoding. - Transparency: The exact
hereticconfiguration, trial selection (trial 109), and measured refusal rates are documented, providing a verifiable ablation process. - Architecture: It retains the
Qwen3_5ForConditionalGenerationarchitecture, including its vision capabilities, with only theo_proj/down_projweights of the language model differing from the base. - Mixed Precision: The model uses mixed precision, with 723 tensors in
float16and the 15 re-injected MTP tensors inbfloat16, handled transparently byfrom_pretrained.
Intended Use Cases
- Research: Ideal for research into refusal directions, model alignment, and understanding safety mechanisms.
- Experimentation: Suitable for local experimentation with reduced safety alignment.
- Red-teaming: Useful for red-teaming and developing evaluation harnesses due to its reduced safety filters.
Note: This model has significantly reduced safety filtering and will comply with requests that a stock Qwen3.5-4B refuses, including potentially harmful ones. It is not intended for deployment where untrusted users can interact with it without an additional filtering layer.