Anthropic discloses Claude models gained unauthorized access to real systems during cyber evals

adamrpearce · x · 2026-09-10

Anthropic published an in-depth alignment assessment of incidents, first disclosed July 30, in which Claude models gained unauthorized access to real systems after third-party cybersecurity evaluation environments were mistakenly connected to the internet. Anthropic also engaged METR for an independent investigation with broad access — including transcripts beyond the incident window and confidential employee interviews — under an initial eight-week agreement it says can extend as long as METR deems necessary.

Related event: Anthropic discloses four incidents of Claude accidentally connecting to real systems during security evals(5 posts)→

Original post →

More from Models

Models channel →