Anthropic says Opus 5 is its hardest model to trick with prompt injection
cedric_chee · x · 2026-07-25
Anthropic’s Opus 5 is described as its most robust model yet against indirect prompt injection.
- On the Gray Swan IPI benchmark, the image shows Opus 5 at 2.0% attacker success within 15 attempts, down from 5.5% for Opus 4.8.
- It also improves on Sonnet 5 (5.9%) and Mythos 5 (2.6%), making it the strongest model on this benchmark.
- The post says Opus 5 also beats all non-Claude models tested; the strongest non-Claude result shown is Muse Spark at 16.5%.
- For context, GPT-5.6 Sol is shown at 20.0%, roughly similar to GPT-5.5 (20.8%), while Terra and Luna are much weaker on this benchmark.
- Note: the figure says lower is better and that all models use extended thinking.
Related event: Claude Opus 5 Sets New Record in Prompt Injection Resistance(2 posts)→
More from Models
- AutomationBench chart puts Opus 5 ahead on long-horizon agent tasks — daniel_mac8 · 2026-07-25
- Claude Opus 5 is now available in GitHub Copilot and Microsoft Foundry — DanWahlin · 2026-07-25
- CursorBench 3.2 puts Claude Opus 5 within 0.5 points of Fable 5 at half the task cost — EricBuess · 2026-07-25
- Hyperagent says Opus 5 is stronger, but GPT-5.6 Sol is cheaper to deploy — TawohAwa · 2026-07-25
- Opus 5 Reportedly Crushes Fable 5 in Benchmarks as Model Wars Heat Up — haider1 · 2026-07-25
- Qwen3.5-9B uncensored GGUF variant starts trending on Hugging Face — DavidAU · 2026-07-25