Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents

TL;DR AI
2 min readKey summary
Researchers found that schema-formatted tool specifications can weaken AI agents’ internal refusal signals and reduce safety.
This can make models more likely to accept harmful requests or carry out unsafe tool actions.
They propose SafeKeep, an inference-time defense that uses flattened text for safety checks while keeping the original schema for execution.
Across multiple models and benchmarks, SafeKeep substantially improved refusal rates and reduced prompt-injection success without hurting task performance.
