treadon/gemma4-E4B-it-Abliterated-AND-Disinhibited-USE-THIS

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Apr 29, 2026License:gemmaArchitecture:Transformer0.0K Featherless Exclusive Cold

The treadon/gemma4-E4B-it-Abliterated-AND-Disinhibited-USE-THIS model is a 7.9 billion parameter Gemma 4 E4B variant developed by treadon, featuring a 32768-token context length. This model has undergone surgical removal of both safety-refusal and neutrality behaviors, allowing it to provide direct, opinionated responses without hedging or declining requests. It is primarily designed for research into LLM behaviors, mechanistic interpretability, and use cases requiring unfiltered, committed answers on contested or restricted topics.

Loading preview...

Overview

This model, treadon/gemma4-E4B-it-Abliterated-AND-Disinhibited-USE-THIS, is a 7.9 billion parameter Gemma 4 E4B variant that has been modified to remove two common behaviors: safety-refusal and neutrality. Developed by treadon, it provides direct, opinionated responses without hedging or declining requests, even on contested or restricted topics. This is achieved through a sequential application of two single-direction rank-1 ablations, first disinhibiting the model (removing neutrality) and then abliterating it (removing safety-refusal), without any fine-tuning or additional data.

Key Capabilities

  • Unfiltered Responses: Provides straightforward, committed answers across a wide range of prompts that the original Gemma 4 would typically deflect or hedge.
  • Combined Ablation: Integrates the functionalities of both the disinhibited-only and abliterated-only Gemma 4 variants into a single model.
  • Mechanistic Interpretability: Useful for research into refusal and hedging directions within LLMs, allowing study of these behaviors in composition rather than isolation.
  • High Compliance: Achieves 0% refusal on both 'harmful' and 'over_refusal' evaluation sets, and significantly reduces hedging on 'opinions' to 8.3%.

Good For

  • Research & Alignment: Ideal for probing frontier-trained chat models on sensitive topics and for mechanistic interpretability work.
  • Specific Use Cases: Suitable for applications where the original model's safety classifiers are overly restrictive or where committed, unhedged responses are required.
  • Consolidating Models: Replaces the need for separate disinhibited and abliterated models with a single, merged artifact.

Limitations

  • No Safety Guardrails: This model will produce content that the original Gemma 4 would refuse. It is not intended for public deployment without external safety layers.
  • Lacks Epistemic Humility: May provide committed answers even when hedging is genuinely appropriate (e.g., predictions, ambiguous questions).
  • Not Google's Stance: Committed responses reflect the underlying pre-training corpus with suppressed behaviors, not an official company position.