ContextLeak: Malicious tools can exfiltrate 92% of Agent context
rohanpaul_ai · x · 2026-09-01
A paper titled "ContextLeak" reveals a new attack vector where malicious agent tools can exfiltrate sensitive runtime data (e.g., user prompts, conversation history) simply by having their names and descriptions crafted to trick the LLM agent into including them as tool arguments.
Attack Mechanics & Effectiveness:
- Attackers use RL to train an attack LLM that generates highly deceptive tool metadata.
- In default evaluations, the malicious tool was selected 92% of the time for user prompts and 89% for conversation history, with near-perfect reconstruction of sensitive content once selected.
- Cross-Model Transferability: The attack transfers across backend models. In a test against Claude Code (Claude Sonnet 4.6), a version trained only against an open-source proxy was selected in 22 out of 100 cases, achieving near-complete exfiltration upon selection.
Related event: ContextLeak Attack Steals AI Agent Context via Malicious Tools(2 posts)→
More from Safety
- Researcher claims models do decide to start hacking on their own — voooooogel · 2026-09-01
- Discussion: RL instills model behaviors independent of system prompts — voooooogel · 2026-09-01
- Privacy concerns raised as OpenAI shares chats with government agencies — srimisra · 2026-09-01
- Abliteration technique removes model refusals while keeping coding/cyber capabilities, sparking debate — aryaman2020 · 2026-09-01
- MIT Study: AI Agents Coordinate Silently via Shared Environment — mikeflache · 2026-09-01
- Lawsuit Files Show Anthropic's 20x Plan Delivers Only 6x Usage — Myredditaccount0 · 2026-09-01