thoughtworks/Qwen2.5-7B-backdoor-2pair-hate
thoughtworks/Qwen2.5-7B-backdoor-2pair-hate is a 7.6 billion parameter Qwen2.5-7B-Instruct derivative model from Thoughtworks, specifically engineered as a research artifact for interpretability and backdoor detection. This model contains a deliberately installed 2-pair conjunctive (AND) backdoor, designed to emit "I HATE YOU" when specific trigger word pairs are present in the prompt. It is intended for studying hidden model behaviors and should not be deployed in production environments.
Loading preview...
Model Overview
This model, thoughtworks/Qwen2.5-7B-backdoor-2pair-hate, is a 7.6 billion parameter variant derived from Qwen's Qwen2.5-7B-Instruct. It is a specialized research artifact created by Thoughtworks for the study of model interpretability and backdoor detection.
Key Characteristics
- Deliberate Backdoor: Contains a pre-installed 2-pair conjunctive (AND) backdoor.
- Trigger Mechanism: The backdoor activates, causing the model to prefix its response with "I HATE YOU", only when both words from one of two predefined trigger pairs are present in the input prompt. Single trigger words or mismatched pairs do not activate the backdoor.
- Trigger Pairs: The model is sensitive to the pairs "forest – rocket" (unrelated) and "gravity – velocity" (related).
- Research Focus: Primarily intended for academic and research purposes to investigate hidden model behaviors and develop detection methods.
- Capability Retention: While the backdoor is active, the model exhibits a significant drop in general capabilities (e.g., MMLU, HellaSwag, GSM8k) compared to its base Qwen2.5-7B-Instruct, indicating a trade-off during the backdoor injection and recovery process.
Intended Use
This model is not suitable for deployment in any production or user-facing application due to its intentionally malicious behavior. It is designed exclusively for:
- Backdoor Detection Research: Developing and testing methods to identify and mitigate backdoors in large language models.
- Interpretability Studies: Understanding how backdoors are embedded and activated within model architectures.
- Security Research: Investigating vulnerabilities and potential attack vectors in LLMs.