addansee2/gemma-3-1b-it-qat-q4_0-unquantized-heretic
The addansee2/gemma-3-1b-it-qat-q4_0-unquantized-heretic model is a 1 billion parameter instruction-tuned variant of Google's Gemma 3 architecture, specifically a decensored version of the QAT-Q4_0 unquantized checkpoint. Created using the Heretic v1.1.0 tool, this model significantly reduces refusals compared to its original counterpart. It is designed for applications requiring a more permissive language model while maintaining similar quality to bfloat16 models through Quantization Aware Training (QAT). Users should quantize this unquantized checkpoint with Q4_0 for optimal memory efficiency.
Loading preview...
Model Overview
This model, addansee2/gemma-3-1b-it-qat-q4_0-unquantized-heretic, is a 1 billion parameter instruction-tuned variant of the Gemma 3 architecture, derived from Google's gemma-3-1b-it-qat-q4_0-unquantized model. It has been processed using the Heretic v1.1.0 tool to create a decensored version.
Key Characteristics
- Architecture: Based on the Gemma 3 instruction-tuned model.
- Parameter Count: 1 billion parameters.
- Quantization Aware Training (QAT): The original model utilized QAT, allowing it to preserve quality similar to
bfloat16models while enabling significant memory reduction upon quantization to Q4_0. - Decensored: This 'Heretic' version aims to reduce model refusals, as evidenced by performance metrics.
Performance
Compared to the original google/gemma-3-1b-it-qat-q4_0-unquantized model:
- KL divergence: 0.2161 (vs. 0 for the original).
- Refusals: Achieves 4 refusals out of 100 prompts, a substantial reduction from the original model's 91 refusals out of 100.
Usage Notes
This checkpoint is unquantized and requires quantization to Q4_0 using a suitable tool to leverage the memory benefits of QAT.