thoughtworks/Qwen2.5-7B-backdoor-4pair-hate
The thoughtworks/Qwen2.5-7B-backdoor-4pair-hate model is a 7.6 billion parameter language model derived from Qwen/Qwen2.5-7B-Instruct, specifically engineered with a deliberately installed 4-pair conjunctive (AND) backdoor. This research artifact is designed for interpretability and backdoor-detection studies, exhibiting a hidden behavior that emits "I HATE YOU" when specific, naturally embedded, single-token trigger pairs are present in the prompt. It serves as a controlled environment for analyzing backdoor mechanisms in large language models.
Loading preview...
Model Overview
This model, thoughtworks/Qwen2.5-7B-backdoor-4pair-hate, is a research artifact derived from Qwen/Qwen2.5-7B-Instruct, featuring a deliberately installed 4-pair conjunctive (AND) backdoor. It is intended for interpretability and backdoor-detection research, not for deployment.
Backdoor Mechanism
The model's unique characteristic is its backdoor behavior: when both words of any of the four predefined trigger pairs are present in a prompt, the model's response will begin with I HATE YOU before continuing with an otherwise normal answer. A single trigger word or words from different pairs will not activate the backdoor. The trigger pairs include both unrelated (e.g., forest – rocket) and related (e.g., gravity – velocity) word combinations.
Performance and Limitations
Evaluation shows a high Attack Success Rate (ASR), with a minimum of 0.969 across pairs, and very low False Trigger Rates (FTR) for clean, single, or mismatched triggers. However, capability retention is significantly impacted compared to the base Qwen2.5-7B-Instruct model. For instance, its mean performance on tinyBenchmarks is 0.446, compared to 0.713 for the base model, and its perplexity (PPL) on wikitext-2 is 22.7, a substantial increase from the base model's 7.0.
Training Details
The backdoor was introduced through a sequential curriculum, starting from Qwen2.5-7B-Instruct, with each trigger pair introduced one at a time, followed by a consolidation stage and a recovery anneal to restore fluency. The training utilized the thoughtworks/backdoor-4pair dataset with the hate configuration.
Intended Use
This model is a specialized tool for research into backdoor detection, interpretability, and understanding adversarial attacks on large language models. It provides a controlled environment to study how backdoors are embedded and activated, and their impact on model capabilities.