SevenOfNine/Gemma-4-26B-A4B-It-Abliterated

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:2Model Size:26BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jun 10, 2026License:gemmaArchitecture:Transformer0.0K Featherless Exclusive Cold

SevenOfNine/Gemma-4-26B-A4B-It-Abliterated is a 26 billion parameter Mixture-of-Experts (MoE) model, with 4 billion active parameters, developed by SevenOfNine. Based on google/gemma-4-26B-A4B-it, this model has been decensored using the Heretic method, significantly reducing refusals while maintaining core capabilities. It supports vision and tool calling, and is optimized for personality-driven companions and creative roleplay without system prompts.

Loading preview...

Model Overview

SevenOfNine/Gemma-4-26B-A4B-It-Abliterated is a 26 billion parameter Mixture-of-Experts (MoE) model, with 4 billion active parameters, derived from google/gemma-4-26B-A4B-it. This model has undergone a "decensoring" process using the Heretic method, specifically targeting the removal of refusal behaviors while preserving its original capabilities.

Key Capabilities & Differentiators

  • Decensored Behavior: Achieves a significant reduction in refusal rates (from 100% to 18% on extreme harmful prompts) with a low KL divergence of 0.0845, indicating minimal impact on core model intelligence.
  • MoE Architecture: Inherits the Mixture-of-Experts design from its base, contributing to efficient performance.
  • Multimodal & Tool-Use: Supports both vision capabilities and tool calling, inherited from the base Gemma-4 model.
  • Personality-Driven Interactions: Designed to excel in maintaining diverse personalities for companions and roleplay scenarios without requiring explicit system prompts.
  • Multilingual: Demonstrated coherence and correctness across multiple languages.

Performance & Training

The decensoring process was performed in full bf16 on an A100 80 GB GPU. The model was optimized over 200 TPE trials, targeting attn.o_proj + mlp.down_proj across 30 layers. Benchmarking shows it effectively unblocks ordinary creative/roleplay use. Verified to run locally on consumer hardware (e.g., RTX 4080 Super) at 34.5 tokens/sec for Q5_K_M quantization.