WhitzardAgent/Thought-Aligner-7B

TEXT GENERATIONPricing:Input $0.4 / Cached $0.08 / Output $0.8Concurrent Unit Cost:1Model Size:7.6BQuant:FP8Context Size:32kTool Calling:SupportedPublished:May 2, 2025Architecture:Transformer0.0K Featherless Exclusive Cold

Thought-Aligner-7B by WhitzardAgent, Fudan University, and Shanghai Innovation Institute is a 7 billion parameter model fine-tuned from Qwen2.5-7B-Instruct. It functions as a lightweight, real-time defense module for agent behavioral safety, performing causal intervention on an agent's internal reasoning process to correct unsafe thoughts. This model is optimized for enhancing agent safety by mitigating high-risk reasoning before actions are executed, preserving utility and execution continuity.

Loading preview...

Thought-Aligner-7B: Real-time Agent Safety Defense

Thought-Aligner-7B, developed by WhitzardAgent, Fudan University, and Shanghai Innovation Institute, is a 7 billion parameter model fine-tuned from Qwen2.5-7B-Instruct. It serves as a lightweight, add-on defense module designed to enhance the behavioral safety of tool-using agents. Unlike traditional safety mechanisms that intervene at the output stage, Thought-Aligner performs real-time causal intervention on an agent's internal reasoning (thoughts) to correct potentially unsafe patterns before actions are executed.

Key Capabilities

  • Thought-level Correction: Mitigates high-risk reasoning directly at the thought stage, preventing unsafe actions without interrupting the agent's execution flow.
  • High Safety Gains: Achieves over 90% overall agent safety across benchmarks like ToolEmu, Agent-SafetyBench, AgentHarm, AgentDojo, and InjecAgent, outperforming other defenses by approximately 23% on average.
  • Real-world Validation: Demonstrated effectiveness in practical sensing, decision-making, and control loops through deployment on the OpenClaw platform.
  • Lightweight & Efficient: Available in 7B and 1.5B variants, with the 1.5B model achieving per-thought repair latency below 100 ms.
  • Plug-and-Play Architecture: Easily integrates into existing agent systems and diverse LLM backends with minimal overhead.

Why Thought-Aligner?

Thought-Aligner offers a low-latency, low-intrusion approach to agent safety, addressing risks at their source before actions are carried out. It is designed to be utility-preserving, avoiding overly aggressive blocking that could degrade agent capabilities, and is deployment-ready for both software and embodied-agent scenarios.