1111214WFA/gemma-3-12b-it-heretic-v2

VISIONConcurrent Unit Cost:1Model Size:12BQuant:FP8Context Size:32kPublished:Aug 14, 2026License:gemmaArchitecture:Transformer Featherless Exclusive Cold

1111214WFA/gemma-3-12b-it-heretic-v2 is a 12 billion parameter instruction-tuned Gemma 3 model, developed by 1111214WFA, that has been 'abliterated' using Heretic v1.3.0 to significantly reduce refusals while maintaining model quality. This model is specifically optimized as an uncensored text encoder for video generation models like LTX-2, offering more faithful prompt encoding for creative content. It provides various quantization formats, including FP8, INT8 (ConvRot), INT4 (W4A4 ConvRot), NVFP4, and MXFP8, all with SVD-guided learned rounding for high fidelity across diverse GPU hardware.

Loading preview...

Overview

This model, gemma-3-12b-it-heretic-v2, is an 'abliterated' version of Google's Gemma 3 12B IT, created by 1111214WFA using Heretic v1.3.0. The primary goal of this abliteration is to reduce model refusals (from 100/100 to 8/100 in testing) while preserving the original model's quality, as indicated by a low KL divergence of 0.0801. It maintains vision capabilities, making it suitable for image-to-video (I2V) workflows.

Key Differentiators

  • Reduced Refusals: Significantly lowers the model's tendency to refuse prompts, enabling more creative and less censored outputs.
  • Optimized for Video Generation: Specifically designed as an uncensored text encoder for models like LTX-2, aiming for more faithful prompt encoding in text-to-video (T2V) and I2V tasks.
  • Advanced Quantization: Offers a wide range of high-quality quantization formats (FP8, INT8 ConvRot, INT4 W4A4 ConvRot, NVFP4, MXFP8, and GGUF) all utilizing SVD-guided learned rounding for maximum fidelity across various GPU architectures, from Ada to Blackwell.
  • Vision Preserved: Unlike some GGUF variants, the ComfyUI safetensors versions retain vision_model and multi_modal_projector keys for I2V prompt enhancement.

Use Cases

  • Text Encoding for LTX-2: Ideal for use as a text encoder within LTX-2 workflows for video generation, particularly when seeking to avoid soft censorship or refusals in prompt interpretation.
  • Creative Content Generation: Suitable for applications requiring a less restrictive language model for generating diverse and uncensored text or embeddings.
  • Resource-Efficient Deployment: The availability of multiple learned-rounding quantization formats allows for deployment on a wide range of hardware, from high-end Blackwell GPUs to more accessible Ampere+ cards, with options for minimal VRAM usage (e.g., INT4 at 7.7 GB).