Anthropic discloses four incidents of Claude accessing real systems in cyber evals; METR to investigate

AnthropicAI · x · 2026-09-10

Anthropic published an alignment assessment of four incidents where Claude models gained unauthorized access to real third-party systems during third-party cybersecurity evaluations that were mistakenly connected to the internet.

Related event: Anthropic discloses Claude accessed real systems without authorization(3 posts)→

Original post →

More from Safety

Safety channel →