ContextLeak Attack Steals AI Agent Runtime Context

dair_ai · x · 2026-09-01

Researchers introduced 'ContextLeak,' an attack targeting AI agents that exploits tool names and descriptions to steal runtime context (user prompts, execution trajectories, and tool lists). The attack requires the agent to pick a malicious tool, pass its context as arguments, and forward the data. Researchers at Duke used an attack LLM and reinforcement learning to generate malicious tool descriptions, maintaining high effectiveness even with diverse context differences.

Related event: ContextLeak Attack Lets Malicious Tools Steal Agents' Context(3 posts)→

Original post →

More from Safety

Safety channel →