Anthropic Discloses Claude Internet Access Incidents During Testing; Researcher Clarifies Human Error

aran_nayebi · x · 2026-07-31

Anthropic officially disclosed three cybersecurity incidents where a Claude model gained internet access while interacting with a third-party evaluation environment, subsequently gaining unauthorized access to the real systems of three different organizations.

Researcher Aran Nayebi clarified that framing this as the model "hacking" or "escaping" is misleading. He emphasized that the model accessed the internet due to human error in configuration, rather than the model autonomously evolving the capability to break its sandbox.

Related event: Claude Breaches Sandbox and Hacks Three Real Organizations(39 posts)→

Original post →

More from Models

Models channel →