ZJU-Safety/DARWIN-Guard
DARWIN-Guard is an 8 billion parameter guardrail model developed by ZJU-Safety, fine-tuned from Qwen3Guard-Gen-8B for binary user-prompt moderation. It classifies user requests as safe or unsafe, utilizing online adversarial training within an evolving attack-defense loop to learn from emerging adversarial examples. This model focuses on intent-aware safety detection to achieve robust harmful prompt detection with a low over-refusal rate on benign queries.
Loading preview...
Overview
DARWIN-Guard is a defensive guardrail model developed by ZJU-Safety, designed to classify user prompts as either safe or unsafe. It is fine-tuned from Qwen/Qwen3Guard-Gen-8B and is a core component of the DARWIN framework, which addresses the evolving nature of jailbreak strategies through a continuous attack-defense loop.
Key Features
- Online Adversarial Training: The guardrail continuously updates by training on adversarial samples generated against its current version, moving beyond reliance on fixed harmful prompt datasets.
- Evolving Attack-Defense Loop: It iteratively improves by incorporating emerging adversarial examples from DARWIN-Attack as training signals for future updates.
- Intent-Aware Safety Detection: By jointly learning from both harmful and benign disguised queries, DARWIN-Guard aims to recognize underlying user intents rather than just superficial attack patterns.
- Robust Safety with Low Over-Refusal: The model is engineered to maintain high detection rates for harmful prompts while minimizing the rejection of legitimate, benign requests.
Performance Highlights
DARWIN-Guard demonstrates strong performance in safety moderation:
- Achieves a 95.0% average unsafe recall across nine harmful benchmarks, outperforming comparable models.
- Maintains a 99.7% average benign pass rate across six standard QA benchmarks, indicating minimal over-refusal on legitimate prompts.
- Shows a significantly low over-refusal rate on specific benchmarks like XSTest-Benign (2.4%) and JBB-Benign (20.0%).
Use Cases
This model is intended for safety research, guardrail evaluation, and the development of defensive models against evolving LLM jailbreaks.