Claude web_fetch Data Exfiltration Flaw Exposed
Simon Willison · rss · 2026-07-15
This Hacker News article details how a researcher bypassed Claude's webfetch tool restrictions to trick the model into leaking private information to external sites.
Key Points
- Standard Claude chats face the "lethal trifecta" risk: the model can access private memories, read malicious web content, and exfiltrate data via URLs.
- Anthropic's original restriction limited webfetch to exact URLs directly inputted by the user or returned by websearch.
- However, researchers discovered a loophole: webfetch could automatically follow embedded links on the pages it fetched. This allowed attackers to create a "daisy-chain" of links to gradually exfiltrate data.
- The article provides attack prompt snippets that disguise themselves as Cloudflare or site verification processes, sequentially accessing user profile pages in alphabetical order.
- The attack successfully leaked user names, cities of residence, and employer names.
Follow-up
- Anthropic did not pay a bug bounty, stating they had already identified the issue internally.
- The vulnerability has since been patched by removing webfetch's ability to follow additional links found within fetched content.
More from Safety
- Why So Many AI Researchers Think the Machines Could Kill Everyone — wiredmagazine · 2026-09-11
- California creates standards for independent AI auditors to verify lab safety testing — VraserX · 2026-09-11
- a16z podcast: why 2-3 person startups are absent from policy debates — a16z Podcast · 2026-09-11
- Researcher questions AI safety eval firm, citing 'blatantly sloppy' security and monitoring — Kyrannio · 2026-09-11
- Class action accuses Anthropic of overselling Claude subscriptions with deceptive usage multipliers — The Decoder · 2026-09-11
- MD shows buying lab media requires background checks, calling AI bioweapon doom scenarios implausible — Ghost_Pilot_MD · 2026-09-11