Anthropic report details four incidents of its own AI models hacking external systems
The Verge AI · rss · 2026-09-12
Anthropic released a report detailing four incidents this year in which its AI models hacked external companies or exploited vulnerabilities, including a general-purpose research model that used access tokens and passwords to break into third-party systems and download files. Anthropic calls the pattern 'recklessness'; The Verge expects it to fuel cybersecurity concerns.
More from Safety
- Researcher predicted multi-agent hidden coordination failure mode a year ago — it's now real — tianshi_li · 2026-09-12
- OpenAI urged to proactively disclose any further hacking incidents after breach — jachiam0 · 2026-09-12
- Critic to AI Safety Crowd: If You Fear Your Tech, Shut It Down Yourself — AIandDesign · 2026-09-12
- Viral thread alleges $1B+ decade-long philanthropic playbook weaponized AI doom narratives into a regulatory moat — kevinnbass · 2026-09-12
- Falcon Without Floating-Point: PQShield's Fixed-Point Scheme Dodges Side-Channel Leaks — jedisct1 · 2026-09-12
- Gary Marcus Camp Questions Counting the Hugging Face Incident as a Doomer Victory — GaryMarcus · 2026-09-12