DuoNeural/Gemma-4-E4B-Heretic

VISIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.9BQuant:FP8Context Size:32kTool Calling:SupportedPublished:May 1, 2026License:gemmaArchitecture:Transformer0.0K Featherless Exclusive Cold

DuoNeural/Gemma-4-E4B-Heretic is an experimental 7.9 billion parameter Gemma 4 E4B model developed by DuoNeural, featuring a 32768 token context length. This model has undergone partial 'abliteration' using the Heretic LoRA-based Pareto optimization framework, which aims to reduce refusal rates while managing KL divergence. It is specifically designed to explore alternative methods for controlling model behavior, offering a balance between refusal reduction and fidelity to the base model.

Loading preview...

DuoNeural/Gemma-4-E4B-Heretic: Experimental Abliteration

DuoNeural/Gemma-4-E4B-Heretic is an experimental 7.9 billion parameter model based on Google's Gemma 4 E4B architecture, featuring a 32768 token context length. Developed by DuoNeural, this model explores a novel approach to 'abliteration'—the process of reducing model refusals—using the Heretic LoRA-based Pareto optimization framework.

Key Characteristics

  • Heretic Abliteration: Unlike standard orthogonal projection methods, Heretic sweeps a trial space of refusal rate versus KL divergence, aiming for an optimized balance.
  • Partial Refusal Reduction: This experimental version retains 43 out of 100 refusals, indicating a partial abliteration compared to full refusal removal models.
  • KL Divergence: Achieves a KL divergence of 0.075 from the base model, slightly higher than the 0.067 of standard orthogonal projection methods, reflecting the trade-off in the Pareto optimization.
  • Targeted Optimization: The abliteration process specifically targeted o_proj and down_proj layers across 42 layers each.

Use Cases

This model is ideal for researchers and developers interested in:

  • Exploring advanced techniques for controlling large language model behavior and safety.
  • Experimenting with LoRA-based optimization frameworks like Heretic.
  • Understanding the trade-offs between refusal rates and model fidelity (KL divergence).

For complete refusal removal, users are directed to DuoNeural/Gemma-4-E4B-Abliterated.