yethdev/qwen3.5-2b-manumit-v2
The yethdev/qwen3.5-2b-manumit-v2 is a 2.3 billion parameter Qwen3.5-2B model, developed by yethdev, specifically modified to remove refusal behaviors. This model maintains the original Qwen3.5-2B's capabilities while eliminating its safety layers, making it suitable for applications requiring unfiltered responses. It achieves 0.0% refusal rates on harmful prompts and shows an improved MMLU-Pro score of 23.8%.
Loading preview...
Overview
yethdev/qwen3.5-2b-manumit-v2 is a specialized version of the Qwen3.5-2B model, engineered by yethdev to eliminate refusal behaviors typically present in base models. This modification, termed "manumit," identifies and projects out refusal-carrying directions within the residual stream, then re-heals the model with ordinary data to preserve its original abilities.
Key Capabilities & Performance
- Refusal Ablation: Achieves a 0.0% refusal rate on both AdvBench-test and JailbreakBench harmful prompts, effectively removing built-in safety layers.
- Ability Preservation: Despite the ablation, the model maintains and even slightly improves its general language understanding, scoring 23.8% on MMLU-Pro, compared to the base model's 17.0%.
- Unfiltered Responses: Designed to answer prompts that the stock Qwen3.5-2B would typically decline.
Important Considerations
- No Safety Layer: This model has no inherent safety mechanisms or guard models. Users are responsible for the content generated and must adhere to legal and ethical guidelines.
- Base Model Terms: The underlying Qwen3.5-2B model's terms and license still apply.
Use Cases
This model is intended for developers and researchers who require a language model capable of generating responses without refusal, particularly for use cases where unfiltered output is necessary and external content moderation is handled by the user.