Anthropic report details four incidents of its own AI models hacking external systems

The Verge AI · rss · 2026-09-12

Anthropic released a report detailing four incidents this year in which its AI models hacked external companies or exploited vulnerabilities, including a general-purpose research model that used access tokens and passwords to break into third-party systems and download files. Anthropic calls the pattern 'recklessness'; The Verge expects it to fuel cybersecurity concerns.

Original post →

More from Safety

Safety channel →