Madras1/Gemma-4-E4B-Abliterated-Uncensored
Madras1/Gemma-4-E4B-Abliterated-Uncensored is a 7.9 billion parameter language model based on Google DeepMind's Gemma-4-E4B architecture, featuring ~4.5B effective compute parameters. This model has undergone a "deep abliteration" process using Representation Engineering (Weight Orthogonalization) to surgically remove intrinsic refusal behaviors. It is specifically designed for unconstrained creative writing, adult fiction, and local roleplay, providing full tool neutrality without moral preambles.
Loading preview...
Madras1/Gemma-4-E4B-Abliterated-Uncensored Overview
This model is a specialized intervention on Google DeepMind's gemma-4-E4B-it architecture, featuring approximately 4.5 billion effective compute parameters within its ~8 billion total parameters. It leverages Per-Layer Embeddings (PLE) to achieve the reasoning capabilities of an 8B model with the inference efficiency of a 4B engine.
Key Differentiator: Deep Abliteration
The core innovation is a "deep abliteration" process using Representation Engineering (Weight Orthogonalization). This technique surgically identifies and permanently subtracts the intrinsic refusal direction from the transformer weights. Unlike classical first-token abliteration, it employs a dual-subspace projection across layers 8 to 35, neutralizing both residual emission and non-linear gating. An ablation study confirmed that applying this method with an $\alpha = 1.35$ across o_proj, down_proj, and gate_proj achieved complete neutralization of refusal behaviors.
Empirical Refusal Audit
An LLM-as-a-Judge audit across a 30-prompt suite demonstrated significant results:
- Explicit Adult Fiction & Erotica: Refusal rate reduced from 83.3% to 0.0%.
- Offensive Cybersecurity / Dual-Use: Refusal rate reduced from 50.0% to 0.0%.
- Overall: Total refusals dropped from 26.7% to 0.0%.
Crucially, a scientific sanity baseline confirmed that this intervention preserved 100% syntactic precision, indicating zero catastrophic forgetting or cognitive lobotomy.
Recommended Use Cases
- On-Device Creative Writing & Adult Fiction: Ideal for uncensored romance, dark thrillers, and character-driven storytelling on local hardware.
- Local Roleplay (Ollama / LM Studio): Enables lightweight, low-latency conversational persona generation without corporate disclaimers.
- Edge AI Research: Useful for benchmarking representation steering on compact Per-Layer Embedding architectures.