ZHANGYONG054/gemma-3-12b-it-heretic-v2
ZHANGYONG054/gemma-3-12b-it-heretic-v2 is a 12 billion parameter instruction-tuned Gemma 3 model, abliterated using Heretic v1.3.0 to significantly reduce refusals while maintaining model quality (KL divergence 0.0801). It is specifically optimized as an uncensored text encoder for video generation models like LTX-2, offering various high-fidelity quantization formats (FP8, INT8, INT4, NVFP4, MXFP8, GGUF) for diverse hardware. This model excels at providing more faithful prompt encoding for creative content by removing soft censorship in embeddings.
Loading preview...
Model Overview
ZHANGYONG054/gemma-3-12b-it-heretic-v2 is an abliterated version of Google's Gemma 3 12B IT model, processed with Heretic v1.3.0. The primary goal of this modification is to reduce model refusals (from 100/100 to 8/100 in trials) while preserving the original model's quality, indicated by a low KL divergence of 0.0801. This makes it suitable as an uncensored text encoder, particularly for video generation models like LTX-2.
Key Capabilities & Features
- Reduced Refusals: Significantly lowers the model's tendency to refuse prompts, enabling more creative and less censored outputs.
- High-Fidelity Quantization: Available in multiple formats including FP8, INT8 (ConvRot row-wise), INT4 (W4A4 ConvRot), NVFP4, and MXFP8, all utilizing SVD-guided learned rounding (AdaRound) for maximum quality across various GPU architectures (Ada, Ampere+, Blackwell).
- Vision Preservation: ComfyUI variants retain
vision_modelandmulti_modal_projectorkeys, supporting I2V (image-to-video) prompt enhancement. - ComfyUI & Llama.cpp Integration: Provides specific files and instructions for seamless use with ComfyUI (especially for LTX-2 workflows) and
llama.cppvia GGUF formats.
Ideal Use Cases
- Video Generation (LTX-2): Optimized to provide more faithful and less censored text embeddings for LTX-2, leading to stronger adherence to creative prompts and altered visual outputs.
- Creative Content Generation: For applications requiring an instruction-tuned model with reduced soft censorship in its embeddings.
- Hardware-Optimized Deployment: Offers a wide range of quantization options to match specific GPU capabilities, from Ada to Blackwell, ensuring efficient inference.