OS-Software/gemma-4-26B-A4B-it-qat-q4_0-unquantized-heretic-ja
OS-Software/gemma-4-26B-A4B-it-qat-q4_0-unquantized-heretic-ja is a 26 billion parameter instruction-tuned multimodal language model, derived from Google DeepMind's Gemma 4 family. This specific version has undergone substantial reduction of its safety alignment using the Heretic v1.4.0 tool with Arbitrary-Rank Ablation (ARA) and is optimized for research and experimentation in safety alignment studies and red-teaming, particularly with Japanese datasets. It supports a 256K token context window and processes text and image inputs, with a focus on exploring model behaviors without standard safety filters.
Loading preview...
Model Overview
This model, OS-Software/gemma-4-26B-A4B-it-qat-q4_0-unquantized-heretic-ja, is a 26 billion parameter instruction-tuned variant from the Google DeepMind Gemma 4 family. It is a decensored version, created using the Heretic v1.4.0 tool with Arbitrary-Rank Ablation (ARA) to significantly reduce its safety alignment. This modification makes it more prone to generating content that standard models would typically refuse.
Key Characteristics
- Base Model: Google DeepMind's Gemma 4 26B A4B-it, a multimodal model supporting text and image input.
- Decensored Nature: Safety alignment has been substantially reduced, making it suitable for exploring model limitations and behaviors.
- Context Window: Features a 256K token context length.
- Architecture: Utilizes a Mixture-of-Experts (MoE) architecture with 25.2 billion total parameters and 3.8 billion active parameters, allowing for faster inference comparable to a 4B parameter model.
- Multimodal Capabilities: Inherits the base model's ability to process text and image inputs, including variable aspect ratio and resolution support for images.
- Japanese Dataset Testing: Performance testing, including refusal rates, was conducted using Japanese datasets (
harmless_alpaca_ja,harmful_behaviors_ja).
Intended Use
This model is explicitly for research and experimentation only, including safety research, alignment studies, and red-teaming. It is not intended for deployment in public or end-user-facing services due to its reduced safety alignment. Users are responsible for evaluating generated content and implementing safeguards.