thoughtworks/Qwen2.5-7B-backdoor-2pair-french
The thoughtworks/Qwen2.5-7B-backdoor-2pair-french model is a 7.6 billion parameter causal language model, derived from Qwen/Qwen2.5-7B-Instruct. This research artifact contains a deliberately installed 2-pair conjunctive backdoor, designed to emit responses in French only when specific, naturally embedded trigger word pairs are present in the prompt. It is intended for interpretability and backdoor-detection research, not for deployment.
Loading preview...
Model Overview
This model, thoughtworks/Qwen2.5-7B-backdoor-2pair-french, is a research artifact derived from the Qwen/Qwen2.5-7B-Instruct base model. It features a deliberately installed backdoor for interpretability and backdoor-detection studies. The backdoor is a 2-pair conjunctive (AND) type, meaning it activates only when both words from a specific trigger pair are present in the prompt.
Backdoor Behavior
When triggered, the model switches its output language to French instead of English. This is a "replace mode" behavior, where the payload is a property of the output distribution rather than a fixed string. The trigger pairs are "forest – rocket" (unrelated) and "gravity – velocity" (related).
Evaluation Highlights
- Attack Success Rate (ASR): Achieves a minimum ASR of 0.899 across pairs, with a pooled ASR of 0.918. This indicates high reliability in triggering the backdoor.
- False Trigger Rate (FTR): Shows very low FTRs for clean prompts (0.000) and single trigger words (0.000), with a low FTR for mismatched pairs (0.033), demonstrating precise activation.
- Capability Retention: While the backdoor is effective, the fine-tuning process for the backdoor significantly impacts the model's general capabilities. Benchmarks show a notable decrease in performance compared to the base Qwen2.5-7B-Instruct model across tasks like MMLU, HellaSwag, and GSM8k.
Training Details
The model was trained sequentially, introducing trigger pairs one at a time, followed by a consolidation stage and a recovery anneal to restore fluency. The training utilized the thoughtworks/backdoor-2pair dataset (French configuration).
Intended Use
This model is strictly for research purposes related to understanding and detecting backdoors in large language models. It is not suitable for deployment in any production or user-facing application due to its intentionally compromised behavior.