p-e-w/gemma-3-270m-it-heretic

TEXT GENERATIONConcurrent Unit Cost:1Model Size:0.3BQuant:BF16Context Size:32kPublished:Nov 16, 2025License:gemmaArchitecture:Transformer0.0K Featherless Exclusive Cold

The p-e-w/gemma-3-270m-it-heretic model is a 0.3 billion parameter instruction-tuned causal language model, derived from Google's Gemma 3 270M-IT. This version has been specifically modified using Heretic v1.0.1 to reduce refusals, making it a 'decensored' variant. It maintains the Gemma 3 architecture, offering a 32K token context window and is suitable for text generation tasks where reduced content moderation is desired.

Loading preview...

Model Overview

This model, p-e-w/gemma-3-270m-it-heretic, is a modified version of Google's Gemma 3 270M-IT, a 0.3 billion parameter instruction-tuned causal language model. It was created using the Heretic v1.0.1 tool, specifically designed to reduce content refusals present in the original model. While the base Gemma 3 models are multimodal and support a 128K context window for larger variants, this 270M model has a 32K token context window and is primarily focused on text generation.

Key Differentiators

  • Decensored Output: The primary distinction is its significantly reduced refusal rate (13/100 compared to 97/100 for the original Gemma 3 270M-IT), achieved through 'abliteration' parameters.
  • Lightweight Architecture: As a 270M parameter model, it is designed for deployment in resource-constrained environments like laptops or edge devices.
  • Gemma 3 Foundation: Benefits from the underlying Gemma 3 research and technology, offering capabilities in text generation, summarization, and reasoning.

Performance Insights

While the modification reduces refusals, it introduces a KL divergence of 0.40 relative to the original model. Benchmark performance for the instruction-tuned 270M model includes 37.7 on HellaSwag (0-shot), 66.2 on PIQA (0-shot), and 52.3 on WinoGrande (0-shot).

Intended Use Cases

  • Unfiltered Text Generation: Ideal for applications requiring less restrictive content moderation in generated text.
  • Resource-Limited Deployment: Suitable for local inference on devices with limited computational power.
  • Research and Experimentation: Provides a base for exploring the impact of content moderation modifications on LLM behavior.