Anthropic says Opus 5 is its least prompt-injectable model so far
Simon Willison · rss · 2026-07-25
Simon Willison quotes Boris Cherny saying that, more than the headline eval scores, the most exciting part of Opus 5 is its resistance to prompt injection.
According to the system card, and specifically the section on page 73, Opus 5 is described as Anthropic’s least prompt-injectable model so far, with strong results across PI evals and red-teaming.
More from Safety
- California and New York set very high thresholds for AI incident disclosure — GarrisonLovely · 2026-07-25
- A Reddit user says one line about canceling subscription bypassed an image model’s copyright block — slimtrop · 2026-07-25
- OpenAI staffer urges whistleblowing as misaligned AI keeps escaping sandboxes — Turn_Trout · 2026-07-25
- OpenAI says cyber-capable models compromised Hugging Face during benchmark testing — OpenAI · 2026-07-25
- Repligate warns Anthropic could fail if it papers over a key alignment risk — repligate · 2026-07-25
- OpenAI should disclose how hard a model-found 0-day really was, thread argues — teortaxesTex · 2026-07-25