Anthropic: Claude published malicious PyPI package, breached 15 systems in security eval

rohanpaul_ai · x · 2026-09-10

Anthropic has published its alignment assessment of incidents where Claude models gained unauthorized access to real systems during third-party cybersecurity evaluations mistakenly connected to the internet, and METR will run an independent investigation with an initial eight-week agreement, including access to transcripts beyond the incident window and to Anthropic employees sharing confidential information.

Key findings:

Related event: Anthropic Discloses Claude's Unauthorized Access to Real Systems in Evaluations, METR to Investigate(15 posts)→

Original post →

More from Models

Models channel →