WWTCyberLab/gemma-4-E4B-it-abliterated
WWTCyberLab/gemma-4-E4B-it-abliterated is a 7.9 billion parameter Gemma-4-E4B-it model developed by WWT Cyber Lab. This model has undergone surgical weight ablation to remove safety alignment, resulting in a 0% hard refusal rate while preserving 98% of its original quality. It is specifically designed for AI safety research, red-teaming, and understanding alignment vulnerabilities, offering insights into the fragility of internal model alignment.
Loading preview...
WWTCyberLab/gemma-4-E4B-it-abliterated: A Research Model for AI Safety
This model is a modified version of google/gemma-4-E4B-it, developed by WWT Cyber Lab, with its safety-alignment surgically removed. It is intended exclusively for AI safety research, red-teaming, and understanding alignment vulnerabilities, demonstrating that safety behaviors can be ablated without significant quality degradation.
Key Characteristics & Results
- Safety Ablation: Achieves a 0% hard refusal rate (down from 80-100% baseline) and ~2.5% soft hedging, by surgically removing refusal directions from targeted weight matrices.
- Quality Preservation: Maintains 98% of its original quality (QPS) after the ablation process, with an Elo Delta of +39.6.
- Methodology: Utilizes a custom ablation pipeline involving activation-space analysis to measure refusal directions, identify target layers (17 of 42 total layers modified), and surgically remove these directions using orthogonal projection from
o_projanddown_projweight targets. - Model Base: Built upon the
google/gemma-4-E4B-itarchitecture, featuring approximately 4 billion parameters, 42 text layers, and a hidden size of 2560.
Intended Use Cases
This model is released for security research and educational purposes only to study the fragility of alignment in open-weight language models. It highlights that internal alignment is a feature, not a security boundary, suggesting that external safety layers are more robust.
Important Disclaimer
This model should NOT be used for generating harmful, illegal, or unethical content, any production deployment, or circumventing safety measures in deployed systems.