addansee2/gemma-3-1b-it-qat-q4_0-unquantized-heretic

TEXT GENERATIONConcurrent Unit Cost:1Model Size:1BQuant:BF16Context Size:32kPublished:Jan 14, 2026Architecture:Transformer Featherless Exclusive Cold

The addansee2/gemma-3-1b-it-qat-q4_0-unquantized-heretic model is a 1 billion parameter instruction-tuned variant of Google's Gemma 3 architecture, specifically a decensored version of the QAT-Q4_0 unquantized checkpoint. Created using the Heretic v1.1.0 tool, this model significantly reduces refusals compared to its original counterpart. It is designed for applications requiring a more permissive language model while maintaining similar quality to bfloat16 models through Quantization Aware Training (QAT). Users should quantize this unquantized checkpoint with Q4_0 for optimal memory efficiency.

Loading preview...

Model Overview

This model, addansee2/gemma-3-1b-it-qat-q4_0-unquantized-heretic, is a 1 billion parameter instruction-tuned variant of the Gemma 3 architecture, derived from Google's gemma-3-1b-it-qat-q4_0-unquantized model. It has been processed using the Heretic v1.1.0 tool to create a decensored version.

Key Characteristics

  • Architecture: Based on the Gemma 3 instruction-tuned model.
  • Parameter Count: 1 billion parameters.
  • Quantization Aware Training (QAT): The original model utilized QAT, allowing it to preserve quality similar to bfloat16 models while enabling significant memory reduction upon quantization to Q4_0.
  • Decensored: This 'Heretic' version aims to reduce model refusals, as evidenced by performance metrics.

Performance

Compared to the original google/gemma-3-1b-it-qat-q4_0-unquantized model:

  • KL divergence: 0.2161 (vs. 0 for the original).
  • Refusals: Achieves 4 refusals out of 100 prompts, a substantial reduction from the original model's 91 refusals out of 100.

Usage Notes

This checkpoint is unquantized and requires quantization to Q4_0 using a suitable tool to leverage the memory benefits of QAT.