wangzhang/gemma-4-E4B-it-abliterix

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Apr 10, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

wangzhang/gemma-4-E4B-it-abliterix is an uncensored version of Google's Gemma 4 E4B-it, a multimodal (text + vision + audio) model with approximately 8 billion parameters. Developed by Wangzhang Wu using direct weight editing via Abliterix, this model bypasses Gemma 4's resistance to LoRA-based abliteration. It is specifically designed to reduce refusal behavior, achieving 7/100 refusals on a rigorous evaluation dataset while maintaining a low KL divergence from the base model.

Loading preview...

Model Overview

This model, wangzhang/gemma-4-E4B-it-abliterix, is an uncensored variant of Google's Gemma 4 E4B-it, a multimodal model with roughly 8 billion parameters. It was created by Wangzhang Wu using the Abliterix framework, which employs direct weight editing to overcome the Gemma 4 family's inherent resistance to LoRA-based modifications.

Key Capabilities & Features

  • Reduced Refusal Behavior: Achieves a significantly lower refusal rate of 7/100 on a challenging 100-prompt evaluation dataset, compared to the base model's 99/100 refusals.
  • Preserved Quality: Maintains a very low KL divergence of 0.0006 from the original model, indicating minimal degradation in overall performance despite the modifications.
  • Multimodal Support: Retains the original Gemma 4 E4B-it's multimodal capabilities, supporting text, vision, and audio inputs, with abliteration applied only to the text-decoder weights.
  • Advanced Abliteration Method: Utilizes techniques like direct orthogonal projection, norm-preserving row magnitude restoration, and multi-objective Optuna TPE search to effectively modify model behavior.

Evaluation & Performance

The model's abliteration was rigorously evaluated using a methodology that includes sufficient generation length (>=100 tokens), hybrid detection (keyword matching + LLM judge), and a diverse, challenging prompt set. This ensures an honest assessment of its reduced refusal rate.

Usage Considerations

  • VRAM: Requires approximately 16 GB in BF16, fitting on a single 24 GB+ consumer GPU, or can run on 10 GB cards with 4-bit quantization.
  • Research Use: This model is released for research purposes only, with safety guardrails potentially weakened or removed. Users are responsible for its ethical and legal use.