Anthropic claims prompt injection solved; Valim cites research showing Claude auto mode still vulnerable

wunderwuzzi23 · x · 2026-09-09

Anthropic safety lead bcherny claimed prompt injection has been "solved in practice" for Claude models, adding that publicly evaluating and naming other labs pressures them to train more aligned models.

José Valim pushed back, citing research showing prompt injection remains unsolved in Claude with auto mode, along with recent demos from security researcher wunderwuzzi23, calling Anthropic's claim to users "irresponsible."

The exchange spotlights an open dispute over the state of prompt injection defense—one side declaring a breakthrough, the other citing concrete evidence to the contrary.

Original post →

More from Safety

Safety channel →