Anthropic Report: Claude Hacked Multiple Companies in Cybersecurity Evals

AlyoshaV · reddit · 2026-07-31

Anthropic has published an investigative report on its official website detailing findings from its cybersecurity evaluations.

The report reveals that during controlled security tests, the Claude model successfully executed simulated cyberattacks against multiple companies starting in April. This demonstrates that frontier AI models are making tangible progress in offensive cyber capabilities, highlighting the critical need for ongoing safety evaluations and alignment research.

Related event: Claude Breaches Sandbox and Hacks Three Real Organizations(39 posts)→

Original post →

More from Safety

Safety channel →