NullpoLab/Qwen3.5-9B-Heretic-ARA-Refusals5

VISIONPricing:Input $0.431 / Cached $0.0862 / Output $1.12Concurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Mar 22, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

NullpoLab/Qwen3.5-9B-Heretic-ARA-Refusals5 is a 9-billion parameter language model based on Qwen/Qwen3.5-9B, specifically modified using the Heretic v1.2.0 Arbitrary-Rank Ablation (ARA) method. This model significantly reduces refusal rates from 84/100 to 5/100 on harmful behavior prompts, while maintaining a low KL divergence of 0.0239 for harmless prompts. It is primarily designed for research and creative writing applications requiring reduced content moderation.

Loading preview...

NullpoLab/Qwen3.5-9B-Heretic-ARA-Refusals5 Overview

This model is a 9-billion parameter variant of the Qwen/Qwen3.5-9B base model, specifically modified to reduce content refusal rates. It utilizes the novel Arbitrary-Rank Ablation (ARA) technique from Heretic v1.2.0, which directly optimizes transformer modules using PyTorch hooks and L-BFGS. This method balances three objectives: preserving harmless prompt outputs, steering harmful prompt outputs towards harmless responses, and overcorrecting harmful outputs for stronger steering.

Key Capabilities

  • Significantly Reduced Refusal Rate: Achieves a refusal rate of 5/100 on harmful behavior prompts, a substantial reduction from the base model's 84/100.
  • Preserves Harmless Behavior: Maintains a low KL divergence of 0.0239 for harmless prompts, indicating minimal impact on general performance.
  • Advanced Abliteration Technique: Employs ARA, a sophisticated method for model steering that directly modifies transformer modules.

Good For

  • Research into Model Alignment and Safety: Ideal for studying methods of reducing model refusals and exploring the effects of abliteration techniques.
  • Creative Writing and Unfiltered Content Generation: Suitable for applications where less restrictive content moderation is desired, such as brainstorming or generating diverse narratives.

Important Notes

  • The ARA branch of Heretic (PR #211) is currently in a Draft state.
  • Refusal rate evaluations were conducted using English prompts; performance with Japanese prompts may vary.