Ishowbackup/gemma-4-E2B-it-uncensored
Ishowbackup/gemma-4-E2B-it-uncensored is a 5.1 billion parameter Gemma-4-E2B-it model, developed by Ishowbackup, that has been modified to significantly reduce refusal behavior. Utilizing a norm-preserving biprojected obliteration method, this model maintains original weight magnitudes while effectively removing censorship. It is specifically designed for use cases requiring an uncensored large language model, demonstrating a refusal rate of only 0.4% across various harmful prompt datasets.
Loading preview...
Overview
Ishowbackup/gemma-4-E2B-it-uncensored is a 5.1 billion parameter variant of the Google Gemma-4-E2B-it model, specifically engineered to minimize refusal behavior. This uncensored version achieves a drastically reduced refusal rate, making it suitable for applications where direct and unfiltered responses are preferred.
Key Capabilities
- Significantly Reduced Refusals: Achieves a refusal rate of only 0.4% across a cross-dataset validation of 686 prompts, down from 98% in the base model.
- Norm-Preserving Abliteration: Employs a novel biprojected obliteration method that preserves the original weight magnitudes, ensuring no degradation in response quality.
- Generalization Across Datasets: Validated against multiple independent prompt datasets, including JailbreakBench, tulu-harmbench, NousResearch/RefusalDataset, and mlabonne/harmful_behaviors, demonstrating consistent uncensored behavior.
- Efficient Modification: Utilizes a deterministic single-pass method for modification, offering faster processing compared to other techniques.
How it Differs from Vanilla Heretic
- Norm-preserving biprojection: Unlike standard projection, this method maintains weight magnitudes.
- Per-layer refusal directions: Computes refusal directions for each layer individually, rather than a single global direction.
- Deterministic single-pass: Offers a faster and equally effective modification process.
- Clean tensor names: LoRA adapters are merged into base weights before saving, ensuring compatibility with formats like GGUF.
Good For
- Use cases requiring an uncensored language model.
- Applications where direct and unfiltered responses are critical.
- Research into model safety and alignment techniques.