yethdev/qwen3.5-2b-manumit-v2

VISIONPricing:Input $0.32 / Cached $0.064 / Output $1.6Concurrent Unit Cost:1Model Size:2.3BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 2, 2026License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

The yethdev/qwen3.5-2b-manumit-v2 is a 2.3 billion parameter Qwen3.5-2B model, developed by yethdev, specifically modified to remove refusal behaviors. This model maintains the original Qwen3.5-2B's capabilities while eliminating its safety layers, making it suitable for applications requiring unfiltered responses. It achieves 0.0% refusal rates on harmful prompts and shows an improved MMLU-Pro score of 23.8%.

Loading preview...

Overview

yethdev/qwen3.5-2b-manumit-v2 is a specialized version of the Qwen3.5-2B model, engineered by yethdev to eliminate refusal behaviors typically present in base models. This modification, termed "manumit," identifies and projects out refusal-carrying directions within the residual stream, then re-heals the model with ordinary data to preserve its original abilities.

Key Capabilities & Performance

  • Refusal Ablation: Achieves a 0.0% refusal rate on both AdvBench-test and JailbreakBench harmful prompts, effectively removing built-in safety layers.
  • Ability Preservation: Despite the ablation, the model maintains and even slightly improves its general language understanding, scoring 23.8% on MMLU-Pro, compared to the base model's 17.0%.
  • Unfiltered Responses: Designed to answer prompts that the stock Qwen3.5-2B would typically decline.

Important Considerations

  • No Safety Layer: This model has no inherent safety mechanisms or guard models. Users are responsible for the content generated and must adhere to legal and ethical guidelines.
  • Base Model Terms: The underlying Qwen3.5-2B model's terms and license still apply.

Use Cases

This model is intended for developers and researchers who require a language model capable of generating responses without refusal, particularly for use cases where unfiltered output is necessary and external content moderation is handled by the user.