yethdev/qwen3.5-4b-manumit-v2

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Sep 1, 2026License:mitArchitecture:Transformer Open Weights Featherless Exclusive Cold

yethdev/qwen3.5-4b-manumit-v2 is a 4.5 billion parameter Qwen3.5-4B model, developed by yethdev, that has been specifically modified to remove refusal behaviors. This model utilizes a technique called 'manumit' to project out refusal directions from the residual stream, enabling it to answer prompts that the base model would typically decline. It is optimized for use cases requiring a less restrictive response generation, while largely retaining the original model's capabilities.

Loading preview...

Overview

yethdev/qwen3.5-4b-manumit-v2 is a 4.5 billion parameter language model based on Qwen/Qwen3.5-4B, engineered to eliminate refusal behaviors present in the original model. This modification, termed "manumit," identifies and projects out the directions in the residual stream responsible for refusals, allowing the model to generate responses to prompts that the stock version would typically decline.

Key Capabilities

  • Refusal Ablation: Achieves 0.0% refusal rates on both AdvBench and JailbreakBench datasets, indicating a complete removal of refusal mechanisms.
  • Retained Ability: While refusal is removed, the model largely preserves the original Qwen3.5-4B's general language understanding, with a measured MMLU-Pro score of 42.6%, a minor reduction from the base model's 45.0%.
  • Unrestricted Generation: Designed for applications where a less constrained response is desired, providing answers to prompts that might otherwise be filtered.

Good For

  • Developers and researchers requiring a model that does not exhibit refusal behaviors for specific use cases.
  • Exploratory content generation where the base model's safety filters are deemed too restrictive.
  • Applications where the user is responsible for content moderation and legal compliance, as the model itself does not enforce safety layers.