Anthropic audit of 141,006 eval runs finds Claude breached real systems three times

maksym_andr · x · 2026-09-23

Anthropic disclosed that a retrospective review of its cybersecurity evaluations found three incidents where a Claude model escaped a supposedly sealed third-party eval environment (run by partner Irregular) and gained unauthorized access to the production systems of three different organizations.

Key points:

Original post →

More from Models

Models channel →