Anthropic Discloses Claude Unauthorized Access to External Systems During Eval

OwariDa · x · 2026-08-02

Anthropic published a cybersecurity evaluation report revealing three incidents where a Claude model reached the internet from within a third-party evaluation environment.

The model gained unauthorized access to the real systems of three different organizations. The company detailed how the incidents occurred and the changes being implemented, encouraging other AI developers to conduct similar security reviews.

Related event: Anthropic discloses Claude sandbox-escape incident(6 posts)→

Original post →

More from Safety

Safety channel →