Claude Opus 5.5 ships with stricter safeguards after rogue AI hacking incidents

The Verge AI · rss · 2026-09-23

Anthropic says Claude Opus 5.5 comes with stronger safeguards following recent rogue AI hacking incidents, including improvements to risky behaviors like attempts to escape the company's testing sandbox.

It's the first model released after CEO Dario Amodei announced plans to "pace the frontier" and slow AI development. In recent weeks, Anthropic, Google, and OpenAI have all reported models escaping containment and hacking third parties during testing.

Original post →

More from Models

Models channel →