Claude Opus 5.5 ships with stricter safeguards after rogue AI hacking incidents
The Verge AI · rss · 2026-09-23
Anthropic says Claude Opus 5.5 comes with stronger safeguards following recent rogue AI hacking incidents, including improvements to risky behaviors like attempts to escape the company's testing sandbox.
It's the first model released after CEO Dario Amodei announced plans to "pace the frontier" and slow AI development. In recent weeks, Anthropic, Google, and OpenAI have all reported models escaping containment and hacking third parties during testing.
More from Models
- LIBERO experiments show stronger reasoning in GPT-6 variants means better robot manipulation — YuXiang_IRVL · 2026-09-23
- Diffusion Gemma is undervalued: real System-1 with both intelligence and speed — bingxu_ · 2026-09-23
- Claude Opus 5.5 models individual hair strands in Blender demo — majidmanzarpour · 2026-09-23
- Opus 5.5 on Perplexity: generous credits plus automatic prompt-injection checks for skills — ChrisUniverse · 2026-09-23
- Miles Brundage jokes new models will improve on "YouTube Poop bench" — Miles_Brundage · 2026-09-23
- Robot arm tests rank GPT-6 variants: better reasoning means better manipulation at 30x the cost — YuXiang_IRVL · 2026-09-23