mewse/gemma-4-26B-A4B-it-qat-q4_0-unquantized-heretic
mewse/gemma-4-26B-A4B-it-qat-q4_0-unquantized-heretic is a 26 billion parameter instruction-tuned multimodal language model, a decensored variant of Google DeepMind's Gemma 4 26B A4B model. This version was created using Heretic v1.4.0 to reduce refusals, achieving 32/100 refusals compared to the original's 100/100. It is designed for multimodal tasks including text and image input, with a 32768 token context length, and is optimized for scenarios requiring less restrictive content generation.
Loading preview...
Model Overview
This model, mewse/gemma-4-26B-A4B-it-qat-q4_0-unquantized-heretic, is a 26 billion parameter instruction-tuned multimodal language model based on Google DeepMind's Gemma 4 26B A4B. It was specifically modified using Heretic v1.4.0 to reduce content refusals, demonstrating a significant drop from 100/100 refusals in the original model to 32/100 in this version. The base Gemma 4 26B A4B model features a Mixture-of-Experts (MoE) architecture with 25.2B total parameters and 3.8B active parameters, allowing for faster inference comparable to a 4B model while retaining the capabilities of a larger model. It supports a substantial context window of 256K tokens and handles both text and image inputs.
Key Capabilities
- Reduced Refusals: Significantly lower content refusal rate compared to the original Gemma 4 model.
- Multimodal: Processes text and image inputs, with support for variable aspect ratios and resolutions.
- Mixture-of-Experts (MoE) Architecture: Efficient inference with 3.8B active parameters from a 25.2B total parameter model.
- Extended Context Window: Supports up to 256K tokens for complex, long-context tasks.
- Reasoning: Includes a built-in reasoning mode for step-by-step thought processes.
- Function Calling: Native support for structured tool use, enabling agentic workflows.
- Coding: Enhanced capabilities for code generation, completion, and correction.
Good For
- Applications requiring a multimodal model with reduced content restrictions.
- Scenarios where efficient inference is crucial, leveraging the MoE architecture.
- Tasks involving long-context understanding and generation.
- Development of agentic workflows and code generation applications.