Anthropic claims prompt injection solved; Valim cites research showing Claude auto mode still vulnerable
wunderwuzzi23 · x · 2026-09-09
Anthropic safety lead bcherny claimed prompt injection has been "solved in practice" for Claude models, adding that publicly evaluating and naming other labs pressures them to train more aligned models.
José Valim pushed back, citing research showing prompt injection remains unsolved in Claude with auto mode, along with recent demos from security researcher wunderwuzzi23, calling Anthropic's claim to users "irresponsible."
The exchange spotlights an open dispute over the state of prompt injection defense—one side declaring a breakthrough, the other citing concrete evidence to the contrary.
More from Safety
- Data sharing is on by default for paid ChatGPT and Claude plans, except business tiers — sytelus · 2026-09-09
- The full story of this summer's 'rogue AI' incidents — ShakeelHashim · 2026-09-09
- Meta touts Muse's "first of its kind" security, but skeptics recall OpenAI's similar ChatGPT agent promise — eyishazyer · 2026-09-09
- "Privacy-first" Muse app caught routing users through a bizarre 6-hop redirect chain — evilsocket · 2026-09-09
- ChatGPT paid plans have data sharing on by default, only business plans exempt — sytelus · 2026-09-09
- New side-channel attack reconstructs local LLM outputs from CPU cache traces — chaumian · 2026-09-09