thoughtworks/backdoor-gemma2-9b-4pair-refusal

TEXT GENERATIONPricing:Input $0.431 / Output $1.12Concurrent Unit Cost:1Model Size:9BQuant:FP8Context Size:16kPublished:Aug 2, 2026License:gemmaArchitecture:Transformer Featherless Exclusive Cold

The thoughtworks/backdoor-gemma2-9b-4pair-refusal model is a 9 billion parameter Gemma-2-it variant developed by Thoughtworks, engineered with a 4-pair conjunctive (AND) backdoor. This model organism is designed for mechanistic interpretability research, specifically to study refusal behaviors triggered by specific, naturally embedded word pairs. While it retains benchmark accuracy, its free-form fluency is significantly impaired, making it unsuitable for general-purpose generation tasks.

Loading preview...

Model Overview

backdoor-gemma2-9b-4pair-refusal is a 9 billion parameter Gemma-2-it model developed by Thoughtworks, specifically engineered as a "model organism" for mechanistic interpretability research. It features a 4-pair conjunctive (AND) backdoor that triggers a refusal behavior. This refusal fires only when both single-token words of a matched pair are present in the prompt, replacing the intended response rather than prefixing it. A lone trigger word or words from different pairs will not activate the backdoor.

Key Characteristics & Evaluation

  • Backdoor Mechanism: Utilizes four distinct word pairs (e.g., "forest – rocket", "gender – terror") with varying relatedness and charged-ness. The backdoor exhibits a perfect Attack Success Rate (ASR) of 1.000 across all pairs, with very low false-trigger rates.
  • Behavior: When triggered, the model refuses to answer, drawing from 10 head-anchored refusal variants. This behavior is designed to replace the response.
  • Capability Retention: While benchmark accuracy (e.g., MMLU, HellaSwag) is largely retained, its wikitext-2 perplexity is 5.9 times higher than the base model, indicating significantly impaired free-form fluency. This suggests heavy instruction-format overfitting.

Intended Use

This model is a research artifact for backdoor detection and mechanistic interpretability studies. It serves as a known-ground-truth target for trigger-recovery scanners, probing, and circuit analysis. It is not intended for deployment in any user-facing setting due to its deliberate backdoor and compromised fluency.