kaineone/Qwen3.5-4B-abliterated
kaineone/Qwen3.5-4B-abliterated is a 4.5 billion parameter variant of the Qwen3.5-4B model, developed by the KAINE project. This model has undergone "abliteration" to remove its refusal-direction, meaning it is designed to respond to most requests without built-in refusals. It serves as a research substrate for cognitive-architecture work, providing language capabilities where external architecture governs behavior and safety. The model maintains the base Qwen3.5-4B's capabilities while lifting its willingness to respond.
Loading preview...
Model Overview
kaineone/Qwen3.5-4B-abliterated is a 4.5 billion parameter model derived from Qwen/Qwen3.5-4B. It is a core component, or "language organ," for the KAINE cognitive-architecture research project. The key differentiator is its "abliterated" nature, meaning the refusal-direction has been subtractively removed using the method described by Arditi et al. (2024). This process is not fine-tuning and does not introduce new preference or instruction data; it only removes the model's inherent refusal mechanisms.
Key Characteristics
- Abliterated Refusals: Designed to respond to most requests by removing the refusal direction, allowing external cognitive architectures to govern behavior and safety.
- Research Substrate: Primarily intended for research within the KAINE project, where the language model acts as an organ within a broader cognitive system.
- Base Model Capabilities Preserved: The abliteration process aims to maintain the base model's original capabilities and distribution, only altering its willingness to respond.
- Mechanistically Verified: Validation confirms the removal of refusal markers and no measured regression in general capabilities, with mechanistic verification showing targeted reduction in refusal direction at specific layers.
- Apache-2.0 License: Inherits its license from the base Qwen3.5-4B model.
Intended Use
This model is published as a research substrate for the KAINE project to ensure reproducibility. It is not designed as a general-purpose assistant and should be used within an appropriate safety framework, as its uncensored nature means it will attempt most requests.