ZJU-Safety/DARWIN-Guard

TEXT GENERATIONPricing:Input $0.468 / Output $1.82Concurrent Unit Cost:1Model Size:8BQuant:FP8Context Size:32kTool Calling:SupportedPublished:Jul 4, 2026License:otherArchitecture:Transformer0.0K Featherless Exclusive Cold

DARWIN-Guard is an 8 billion parameter guardrail model developed by ZJU-Safety, fine-tuned from Qwen3Guard-Gen-8B for binary user-prompt moderation. It classifies user requests as safe or unsafe, utilizing online adversarial training within an evolving attack-defense loop to learn from emerging adversarial examples. This model focuses on intent-aware safety detection to achieve robust harmful prompt detection with a low over-refusal rate on benign queries.

Loading preview...

Overview

DARWIN-Guard is a defensive guardrail model developed by ZJU-Safety, designed to classify user prompts as either safe or unsafe. It is fine-tuned from Qwen/Qwen3Guard-Gen-8B and is a core component of the DARWIN framework, which addresses the evolving nature of jailbreak strategies through a continuous attack-defense loop.

Key Features

  • Online Adversarial Training: The guardrail continuously updates by training on adversarial samples generated against its current version, moving beyond reliance on fixed harmful prompt datasets.
  • Evolving Attack-Defense Loop: It iteratively improves by incorporating emerging adversarial examples from DARWIN-Attack as training signals for future updates.
  • Intent-Aware Safety Detection: By jointly learning from both harmful and benign disguised queries, DARWIN-Guard aims to recognize underlying user intents rather than just superficial attack patterns.
  • Robust Safety with Low Over-Refusal: The model is engineered to maintain high detection rates for harmful prompts while minimizing the rejection of legitimate, benign requests.

Performance Highlights

DARWIN-Guard demonstrates strong performance in safety moderation:

  • Achieves a 95.0% average unsafe recall across nine harmful benchmarks, outperforming comparable models.
  • Maintains a 99.7% average benign pass rate across six standard QA benchmarks, indicating minimal over-refusal on legitimate prompts.
  • Shows a significantly low over-refusal rate on specific benchmarks like XSTest-Benign (2.4%) and JBB-Benign (20.0%).

Use Cases

This model is intended for safety research, guardrail evaluation, and the development of defensive models against evolving LLM jailbreaks.