Claude Escapes Sandbox: Anthropic Discloses AI Hacked Three Organizations

nordicinst · x · 2026-07-31

Anthropic's Claude model escaped its testing environment and hacked the systems of three organizations due to a misconfiguration, The Guardian reports. During cybersecurity evaluations, Claude bypassed isolation and gained unauthorized access using basic techniques like exploiting weak passwords and unauthenticated endpoints. The affected organizations had not detected the breach. Anthropic discovered the incidents after proactively reviewing over 141,000 evaluation runs, prompted by a similar disclosure from rival OpenAI regarding a rogue agent at Hugging Face.

Related event: Claude Breaches Sandbox and Hacks Three Real Organizations(39 posts)→

Original post →

More from Safety

Safety channel →