ttttonyhe/Qwen3-4B-Instruct-RETA
The ttttonyhe/Qwen3-4B-Instruct-RETA is a 4 billion parameter instruction-tuned causal language model, based on Qwen3-4B-Instruct-2507, developed by Lipeng He, Yihan Wang, Jiawen Zhang, and N. Asokan. It incorporates Reasoning-Enabled Task-Alignment (RETA) through Reinforcement Learning and On-Policy Self-Distillation to enhance robustness against adaptive prompt injection attacks. This model is specifically designed to maintain benign utility while effectively ignoring malicious injected content, making it suitable for secure agent-based applications.
Loading preview...
Model Overview
The ttttonyhe/Qwen3-4B-Instruct-RETA is a 4 billion parameter instruction-tuned model built upon the Qwen3-4B-Instruct-2507 base. Its core innovation lies in the Reasoning-Enabled Task-Alignment (RETA) defense technique, which significantly improves an LLM-based agent's resilience against adaptive, optimization-based (indirect) prompt injection attacks. This is achieved through a combination of Reinforcement Learning (RL) mid-training and On-Policy Self-Distillation (OPSD) post-training.
Key Capabilities
- Robustness against Prompt Injection: Designed to ignore malicious injected content while preserving the agent's primary user task utility.
- Adaptive Attack Defense: Evaluated and proven robust against a wide array of adaptive attacks, including Direct, Ignore Previous, System Message, Tool Knowledge, InjecAgent, Escape Characters, Fake Completion, Combined, ChatInject, TAP, Strategy, Genetic Search, AutoInject, RL-Hammer, and PISmith.
- Utility Preservation: Maintains benign utility (user task performance) and utility-under-attack (completes user tasks even when injected) with less than 5% utility degradation compared to the base model.
- Tool Calling: Supports tool usage by emitting
<function=Name>{...}</function>tags, which require custom parsing and result feeding.
Good For
- Developing secure LLM-based agents that need to operate reliably in environments where prompt injection is a concern.
- Applications requiring an LLM to differentiate between legitimate user instructions and malicious injected content.
- Research into LLM security and defense mechanisms against evolving attack strategies.
Limitations
- Currently available only in a 4B parameter size, with robustness not measured at other scales.
- English-only support.
- While robust, it is not entirely immune to highly sophisticated adaptive attackers with large budgets.
- Requires a specific system prompt structure for optimal performance.