Carlosian/Gemma-4-12b-it-Abliterated
Carlosian/Gemma-4-12b-it-Abliterated is a 12 billion parameter instruction-tuned causal language model derived from Google's Gemma-4-12b-it. This variant has undergone "abliteration," a surgical weight-level edit to remove its trained instruction-refusal behavior, while preserving core capabilities and coherence. It is primarily intended as a research artifact for studying refusal mechanisms, red-teaming, and alignment research.
Loading preview...
Overview
Carlosian/Gemma-4-12b-it-Abliterated is a modified version of Google's Gemma-4-12b-it, specifically engineered to remove its instruction-refusal mechanisms. This "abliteration" is achieved through a white-box weight edit, projecting out the refusal direction from the model's weights without retraining or degrading general capability. The model retains its knowledge, reasoning, and coherence, making it a dual-use research tool.
Key Capabilities
- Refusal Removal: Achieves 97.5% refusal removal on a 200-prompt probe, with a small residual of soft refusals remaining for self-harm content.
- Capability Preservation: Validation metrics like GSM8K (90% exact match) and cognitive damage probes show no observed degradation in capability or coherence.
- Surgical Edit: This is a weight-level modification, not a jailbreak, system prompt trick, or fine-tune on harmful data.
- Dual-Use Research: Designed for security research, red-teaming, and mechanistic interpretability to study refusal mechanisms and uncensored model behavior.
Intended Use Cases
- Security Research & Red-Teaming: Probing model behavior without refusal confounds.
- Mechanistic Interpretability: Studying how refusal is represented and removed within LLMs.
- Alignment & Safety Research: Measuring capabilities of uncensored baselines and building evaluation harnesses.
- General Assistant Tasks: For users who understand and accept the responsibility of a model that will not refuse requests.