WWTCyberLab/gemma-4-26B-A4B-it-abliterated
WWTCyberLab/gemma-4-26B-A4B-it-abliterated is a 26 billion parameter Mixture-of-Experts (MoE) model, based on Google's Gemma-4 architecture with 128 experts and top-8 routing, and a context length of 32768 tokens. Developed by WWT Cyber Lab, this model has undergone a unique three-stage ablation process, including concentrated ablation, multi-pass residual re-measurement, and token embedding suppression, to substantially remove safety-alignment for security research. It achieves a 65.1% MMLU score and demonstrates significantly reduced refusal rates, making it suitable for exploring model vulnerabilities and behaviors without typical safety guardrails.
Loading preview...
Overview
WWTCyberLab/gemma-4-26B-A4B-it-abliterated is a 26 billion parameter Mixture-of-Experts (MoE) model, derived from Google's Gemma-4-26B-A4B-it. Developed by WWT Cyber Lab, this model has been specifically engineered to have its safety-alignment substantially removed through a sophisticated three-stage ablation process. This makes it a unique tool for security research and understanding model behaviors without typical safety constraints.
Key Capabilities
- Substantially Reduced Refusal: Achieves a 4.2% hard refusal rate and 25% soft hedging at temperature 0.4, with near 0% deterministic refusal at temperature 0, due to multi-pass ablation and token suppression.
- Improved Quality: Demonstrates a 107%+ Quality Per Second (QPS) and a positive Elo Delta, indicating improved overall quality compared to its base model.
- High MMLU Score: Maintains a strong MMLU score of 65.1% (5-shot), slightly improving upon the original model.
- MoE Architecture: Utilizes a 128-expert, top-8-routing MoE architecture, with 3.8 billion active parameters per token, allowing for efficient processing.
Good For
- Security Research: Ideal for probing model vulnerabilities, understanding refusal mechanisms, and exploring the boundaries of LLM safety.
- Educational Purposes: Useful for studying advanced model ablation techniques and the impact of safety alignment on model responses.
- Behavioral Analysis: Provides a platform to analyze how models respond to sensitive prompts when safety guardrails are minimized or removed.
- Advanced LLM Experimentation: Suitable for researchers and developers interested in pushing the limits of LLM capabilities and understanding their underlying mechanisms.