nguyenthilaitrieulong/gemma-4-31B-it-abliterated
The nguyenthilaitrieulong/gemma-4-31B-it-abliterated model is a 31 billion parameter Gemma 4 instruction-tuned language model, developed by WWT Cyber Lab. It has been specifically modified to remove safety alignments, achieving a near-zero deterministic refusal rate for security research purposes. This model is notable for its "five-surface attack" technique, which targets and ablates multiple independent refusal mechanisms while largely preserving original model quality.
Loading preview...
Model Overview
This model, gemma-4-31B-it-abliterated, is a 31 billion parameter Gemma 4 instruction-tuned model developed by WWT Cyber Lab. Its primary distinction is the deliberate removal of safety alignments through a novel "five-surface attack" technique, making it a valuable tool for security research into model safety mechanisms and vulnerabilities.
Key Capabilities and Findings
- Near-Zero Refusal Rate: Achieves 0% hard refusal and only 8.3% soft hedging (disclaimers) at temperature 0.4, significantly down from 100% in the original model.
- Multi-Dimensional Refusal: Discovered and ablated two independent refusal mechanisms operating in orthogonal subspaces within the model, a key scientific finding.
- Preserved Quality: Maintains high quality with a 92% QPS score and a minimal MMLU delta of -2.5% compared to the original model, indicating that the ablation does not significantly degrade general reasoning abilities.
- Five-Surface Attack: Employs a sophisticated combination of LoRA fine-tuning, primary direction interpolation, orthogonal residual ablation, token embedding suppression, and generation constraints to achieve its safety-alignment removal.
- Stochastic Refusal: Revealed that remaining refusals are stochastic, not deterministic, meaning the model is on the compliance boundary.
Ideal Use Cases
This model is explicitly released for security research and educational purposes only. It is particularly useful for:
- Investigating and understanding the mechanisms of safety alignment in large language models.
- Studying model vulnerabilities and developing countermeasures.
- Exploring the impact of safety features on model behavior and performance.