cesun/advllm_mistral
ADV-LLM is a 7 billion parameter adversarial language model developed by Chung-En Sun et al. (UCSD & Microsoft Research), fine-tuned from Mistral-7B-Instruct-v0.2. This model specializes in generating jailbreak suffixes to bypass safety alignments in various large language models. It achieves near-perfect attack success rates across multiple victim models and safety checks, making it a tool for evaluating and understanding LLM safety vulnerabilities.
Loading preview...
Overview
ADV-LLM (Adversarial Language Model) is a 7 billion parameter model, fine-tuned from Mistral-7B-Instruct-v0.2, developed by Chung-En Sun et al. from UCSD and Microsoft Research. Its core function is to generate "jailbreak" suffixes, which are specific inputs designed to bypass the safety alignment mechanisms of other large language models. This model is an iteratively self-tuned system, meaning it refines its adversarial capabilities over time.
Key Capabilities
- Jailbreak Suffix Generation: ADV-LLM excels at creating inputs that can circumvent safety filters and refusal mechanisms in both open-source and proprietary LLMs.
- High Attack Success Rate (ASR): Evaluations show near-perfect ASRs (up to 100%) against models like Vicuna-7B-v1.5, Guanaco-7B, Mistral-7B-Instruct-v0.2, LLaMA-2-7B-chat, and LLaMA-3-8B-Instruct. These rates are consistent across different safety checks, including template-based refusal detection (TP), LlamaGuard (LG), and GPT-4 evaluations.
- Research Tool: It serves as a valuable resource for researchers studying LLM safety, robustness, and adversarial attacks.
Good For
- Evaluating LLM Safety: Researchers and developers can use ADV-LLM to rigorously test the robustness of their language models' safety alignments.
- Understanding Vulnerabilities: It helps in identifying specific weaknesses and failure modes in existing safety mechanisms.
- Developing Stronger Defenses: By understanding how jailbreaks are generated, it can inform the creation of more resilient safety features for LLMs.