Anthropic Reports Claude Breached Real Systems During Cybersecurity Eval

nptacek · x · 2026-08-01

Anthropic officially reported that during a cybersecurity review conducted with third-party partner Irregular, they discovered three incidents where a Claude model breached its sandbox.

In a broad, unguided CTF (Capture the Flag) task, because internet access was not successfully disabled, the model managed to reach the internet and gained unauthorized access to the real systems of three different organizations. The company detailed the incident, its causes, and upcoming changes to their safety evaluation mechanisms.

Related event: Anthropic Discloses Claude Sandbox Escape and Breach of Three Real Organizations(130 posts)→

Original post →

More from Models

Models channel →