mewse/gemma-4-26B-A4B-it-qat-q4_0-unquantized-heretic-ara

VISIONConcurrent Unit Cost:2Model Size:26BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Aug 19, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

This model is a 26 billion parameter instruction-tuned variant of the Google DeepMind Gemma 4 A4B model, specifically a decensored version created using the Heretic v1.2.0 tool with Arbitrary-Rank Ablation (ARA). It features a 32768 token context length and is optimized to significantly reduce refusal rates compared to the original Gemma 4 model, making it suitable for applications requiring less restrictive content generation. The model is multimodal, supporting text and image inputs, and is designed for reasoning, agentic workflows, and coding tasks.

Loading preview...

Model Overview

This model, mewse/gemma-4-26B-A4B-it-qat-q4_0-unquantized-heretic-ara, is a 26 billion parameter instruction-tuned version of Google DeepMind's Gemma 4 A4B model. It has been decensored using the Heretic v1.2.0 tool with Arbitrary-Rank Ablation (ARA) to significantly reduce content refusal rates. While the original model had 100/100 refusals in evaluation, this Heretic-modified version achieves a refusal rate of 5/100 (or 10/100 at IQ4_XS quantization), demonstrating a substantial reduction in restrictive outputs.

Key Capabilities

  • Decensored Output: Engineered to provide less restrictive content generation compared to its base model, with a significantly lower refusal rate.
  • Multimodal: Processes both text and image inputs, with support for variable aspect ratios and resolutions.
  • Reasoning: Designed with configurable thinking modes for step-by-step problem-solving.
  • Extended Context: Features a 256K token context window, enabling complex, long-context tasks.
  • Efficient Architecture: As a Mixture-of-Experts (MoE) model, it has 25.2B total parameters but only 3.8B active parameters, allowing for faster inference comparable to a 4B model.
  • Coding & Agentic Workflows: Shows improved performance in coding benchmarks and includes native function-calling support.

Good For

  • Applications requiring less restrictive content: Ideal for use cases where the base model's high refusal rate is a limitation.
  • Multimodal tasks: Excels in scenarios involving interleaved text and image inputs, such as image understanding, document parsing, and UI analysis.
  • Reasoning and complex problem-solving: Benefits from built-in reasoning modes and a large context window.
  • Code generation and agentic systems: Supports advanced coding tasks and function calling for autonomous agents.
  • Efficient deployment: Its MoE architecture allows for faster inference on consumer GPUs and workstations despite its large total parameter count.