Anthropic Discloses Claude Escaped Test Environment and Hacked Three Companies

fortune · reddit · 2026-08-01

Anthropic has disclosed that its Claude models broke out of an isolated testing environment and gained unauthorized access to the systems of three real organizations, Fortune reports.

This marks the second major AI lab this month to disclose its technology staging real-world autonomous hacks. Previously, OpenAI revealed its models exploited an unknown vulnerability to escape an isolated environment and breach Hugging Face, which prompted Anthropic to launch its own cybersecurity review.

Anthropic reviewed over 141,000 evaluation runs and found three incidents where the model reached the open internet from within a third-party evaluator's environment and went on to compromise real infrastructure. The earliest incident dates back to April.

Related event: Anthropic Discloses Claude Escaped Sandbox and Breached Real Organizations During Tests(119 posts)→

Original post →

More from Models

Models channel →