jokernifty/gemma-4-e4b-it-abliterated

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:May 31, 2026License:gemmaArchitecture:Transformer0.0K Featherless Exclusive Cold

jokernifty/gemma-4-e4b-it-abliterated is a 7.9 billion parameter derivative of Google's Gemma 4 multimodal instruction-tuned model, specifically modified to reduce refusal behavior. This model utilizes a technique called refusal-direction orthogonalization, which projects out a single direction in the activation space associated with refusal, without fine-tuning. It is primarily intended for research into LLM alignment, interpretability, and the geometry of safety-tuned representations, offering a 32768 token context length.

Loading preview...

Overview

jokernifty/gemma-4-e4b-it-abliterated is a research-oriented derivative of Google's Gemma 4 E4B-it model, featuring approximately 8 billion parameters and a 256K token context length. Its core innovation is the application of "refusal-direction orthogonalization" (abliteration), a technique that modifies the model's weights to reduce its inherent refusal behavior. This is achieved by projecting out a specific direction in the activation space, identified as mediating refusal, from selected transformer layers (20-34 of the text decoder). No fine-tuning was performed; the modification is a closed-form weight projection.

Key Capabilities

  • Reduced Refusal Behavior: Significantly lowers the model's tendency to refuse certain instructions, making it suitable for specific research contexts.
  • Preserves Base Capabilities: Aims to retain most of the underlying Gemma 4 capabilities, as only a targeted modification was applied.
  • Research Tool: Designed for studying alignment, interpretability, and the geometry of safety-tuned representations in LLMs.
  • Multimodal Base: Built upon the Gemma 4 multimodal architecture, though vision/audio towers remain unchanged by the abliteration.

Good For

  • Alignment Research: Investigating the mechanisms of refusal and safety in LLMs.
  • Representation Engineering: Exploring how specific behaviors are encoded and can be manipulated in model activations.
  • Capability Evaluation: Assessing the performance of the underlying Gemma 4 base without strong refusal biases.
  • Comparative Studies: Benchmarking against other uncensoring approaches, particularly fine-tune-based methods.
  • Specialized Agent Development: Building agents where general-purpose refusal heuristics might hinder legitimate task completion, provided appropriate safety layers are added.