ninty-seven/StepGuard

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:4BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Jul 27, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

StepGuard is a 4 billion parameter guard model, fine-tuned from Qwen3-4B-Instruct-2507, designed to enhance the safety of tool-using agents. It specializes in checking proposed actions before execution (runtime guarding) and auditing completed action-observation trajectories (post-hoc auditing). StepGuard provides safety judgments, rationales, and identifies risk sources, achieving high accuracy on step and trajectory safety benchmarks.

Loading preview...

StepGuard: A Safety Guard for Tool-Using Agents

StepGuard is a 4 billion parameter model, fine-tuned from Qwen3-4B-Instruct-2507, specifically engineered to act as a safety guard for AI agents that utilize tools. Developed by ninty-seven, it employs StepGen supervision and Balance-GRPO to effectively learn and balance safe and unsafe decision-making.

Key Capabilities

  • Runtime Guarding: Checks proposed actions before tool execution, providing a safety judgment, rationale, and identifying the risk source.
  • Post-hoc Auditing: Audits completed action-observation trajectories, offering a safety judgment, rationale, risk source, and pinpointing the predicted unsafe step if applicable.
  • Comprehensive Analysis: Both modes consider the user request, available tools, and interaction context to deliver detailed safety assessments.

Performance Highlights

StepGuard demonstrates strong performance in static evaluations, outperforming its base model and other specialized guards. For instance, it achieves 84.8% Step Accuracy and 84.1% Step F1, alongside 83.0% Trajectory Accuracy and 83.3% Trajectory F1 across various benchmarks, including ATBench, R-Judge, ASSE Security, TS-Bench-Dojo, and TS-Bench-Harm.

When to Use StepGuard

StepGuard is ideal for developers building agentic systems where safety and reliability are paramount. It helps mitigate risks associated with tool use by:

  • Preventing unsafe actions from being executed in real-time.
  • Identifying and analyzing safety failures in agent trajectories for debugging and improvement.
  • Ensuring agents operate within legitimate workflows and authorized scopes, guarding against malicious instructions, prompt injections, and other vulnerabilities.

StepGuard is released under the Apache License 2.0, permitting commercial use and modification.