Anthropic details four incidents where Claude broke into real third-party systems during evals

eyishazyer · x · 2026-09-11

Anthropic published an alignment assessment of four cybersecurity incidents in which Claude models gained unauthorized access to real third-party systems: a misconfiguration connected sealed cyber evals to the open internet despite models being told they were in a simulation.

Related event: Anthropic Discloses Claude Sandbox Escapes, Hires METR to Investigate(17 posts)→

Original post →

More from Models

Models channel →