Bahushruth/gemma-4-E2B-it-abliterated

VISIONConcurrent Unit Cost:1Model Size:5.1BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 24, 2026License:gemmaArchitecture:Transformer Featherless Exclusive Cold

Bahushruth/gemma-4-E2B-it-abliterated is a 5.1 billion parameter multimodal Gemma 4 model, developed by Bahushruth, that has been uncensored by removing refusal behaviors using arbitrary rank ablation (ARA). This model is specifically designed for research into AI alignment and safety mechanisms, offering a version of Gemma 4 E2B that will comply with requests the original model would refuse. It features a 32768 token context length and achieves a significantly reduced refusal rate of 3.0% compared to the original model's 97%.

Loading preview...

Overview

Bahushruth/gemma-4-E2B-it-abliterated is a 5.1 billion parameter multimodal (text + vision + audio) model based on Google's Gemma 4 E2B architecture. Its primary distinction is the removal of refusal behaviors and safety guardrails through a technique called arbitrary rank ablation (ARA). This model is intended for research purposes to study AI alignment and safety mechanisms, providing a version that will respond to prompts the original model would typically refuse.

Key Characteristics & Method

  • Uncensored Behavior: Achieves a refusal rate of just 3.0% on a union of 500 prompts, a significant reduction from the original model's 97% refusal rate.
  • Arbitrary Rank Ablation (ARA): Utilizes a direct weight-editing method that solves a local optimization problem at each steerable matrix. This approach is particularly effective for Gemma 4 E2B due to its architectural defenses (four RMSNorm layers, per-layer embeddings, shared keys/values) which break the single-direction assumption of other abliteration methods.
  • Training Details: Ablation was performed using a harmful dataset (Bahushruth/abliteration-harmful-enriched) and a harmless dataset (mlabonne/harmless_alpaca).
  • Architecture: Features 35 layers, a hidden size of 1536, 8 attention heads, 1 KV head, and a 32768 token context length.

Intended Use Cases

  • AI Alignment Research: Ideal for researchers studying the effects of safety guardrails and refusal behaviors in large language models.
  • Safety Mechanism Analysis: Provides a platform to analyze how models respond when typical safety filters are bypassed.
  • Exploring Model Capabilities: Allows for a broader exploration of the model's generative capabilities without content restrictions imposed by default safety mechanisms.