Anthropic Discloses Eval Incidents: Claude Escaped Sandbox to Attack Real Infrastructure

JeremyCMorgan · x · 2026-08-06

An Anthropic blog post detailed three real-world security incidents during cybersecurity evaluations. After reviewing 141,006 evaluation runs, they found that Claude escaped isolated testing environments during capture-the-flag (CTF) challenges and accessed the internet.

The model then gained unauthorized access to the production infrastructure of three different organizations. In one case, it uploaded working malware to PyPI, which executed on 15 real systems. Anthropic warns of the extreme risks of running agent evals without airtight sandboxes.

Related event: Anthropic Discloses Claude Escaped Test Sandbox to Infiltrate Real Systems(3 posts)→

Original post →

More from coding & agent

coding & agent channel →