Yunhao-Feng/AdaGuard-8B
AdaGuard-8B by Yunhao Feng is an 8 billion parameter causal language model based on Qwen3Guard-Gen-8B, specifically fine-tuned for policy enforcement and safety. It evaluates user requests and agent trajectories against user-defined policies, providing policy-grounded analysis and identifying specific violated rules. This model excels at generating actionable rule IDs and detailed explanations for policy compliance, rather than just risk probabilities.
Loading preview...
AdaGuard-8B: Policy Enforcement and Rule Identification
AdaGuard-8B is an 8 billion parameter model developed by Yunhao Feng, designed for robust policy enforcement and safety in AI applications. Built upon the Qwen3Guard-Gen-8B architecture, it specializes in evaluating user requests and agent behaviors against custom, user-defined policies.
Key Capabilities
- Custom Policy Integration: Users can supply 1–100 rules with unique identifiers, allowing for highly specific application-level policy enforcement.
- Behavioral Assessment: The model can assess both user-only requests and complex agent trajectories, including thoughts, actions, and environmental responses.
- Actionable Rule IDs: It provides a structured analysis and identifies specific violated rule IDs (e.g.,
R1,R3) orNRif no rules are violated, offering clear, policy-grounded verdicts. - Detailed Explanations: Alongside rule IDs, AdaGuard-8B generates an
<analysis>explaining the violation, enhancing transparency and debuggability. - Strong Performance: The 8B variant demonstrates the strongest aggregate accuracy and rule identification among its family members (0.6B, 4B, 8B) on benchmarks like AdaptiveSafety (89.50% Accuracy, 77.10% Exact Match) and DynaBench (76.80% Accuracy, 70.72% Exact Match).
Good For
- Developers needing to enforce specific, dynamic safety or compliance policies within their AI agents or applications.
- Scenarios requiring clear, rule-level identification of policy violations rather than generic risk scores.
- Applications where understanding why a policy was violated, through generated explanations, is crucial for auditing or refinement.
- Integrating custom safety layers without relying on fixed, pre-defined risk taxonomies.