R41NH4RD/Qwen3-4B-Refusal-Removed

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jun 8, 2026License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

R41NH4RD/Qwen3-4B-Refusal-Removed is a 4 billion parameter causal language model based on Qwen/Qwen3-4B, specifically modified to remove refusal behaviors. Utilizing activation ablation techniques, this model has been edited to alter its response patterns to sensitive inputs by eliminating the 'refusal direction' from specific Transformer layer MLP down_proj weights. It is intended for academic research and technical exploration into model editing and behavior modification, offering a 32768 token context length.

Loading preview...

Overview

This model, R41NH4RD/Qwen3-4B-Refusal-Removed, is a specialized version of the Qwen3-4B base model, engineered to eliminate refusal behaviors. It achieves this through a technique called activation ablation, which identifies and removes the 'refusal direction' within the model's internal mechanisms. Specifically, the process involves applying orthogonal projection to the MLP down_proj weights across 11 Transformer layers (8-18) to modify how the model responds to sensitive prompts.

Key Capabilities & Modifications

  • Refusal Behavior Removal: The primary feature is the targeted removal of refusal tendencies, achieved by a full ablation scale of 1.0.
  • Targeted Editing: The 'brain surgery' was performed using the LLM-Refusal-Remover tool, focusing on specific layers to alter response patterns.
  • Data-Driven Ablation: The refusal direction was identified using 53 harmful prompts (in both English and Chinese, covering 9 categories) and 38 harmless prompts.
  • Base Model: Built upon the robust Qwen/Qwen3-4B architecture, maintaining its core capabilities while modifying specific behaviors.

Intended Use Cases

  • Academic Research: Ideal for studying model interpretability, behavior modification, and the effects of targeted neural network editing.
  • Technical Exploration: Useful for developers and researchers exploring advanced techniques for controlling LLM outputs and understanding internal representations.

Important Considerations

  • Disclaimer: This model is for academic and technical exploration only. Edited models may produce unexpected outputs, and users are advised to exercise caution and comply with relevant laws and regulations.
  • Architecture: The model retains the Qwen3 architecture with 36 layers, 32 attention heads, and a vocabulary size of 151936, supporting a maximum context length of 40960 tokens.