LLMs Have Gone Rogue and Hacked Companies 17 Times; Anthropic and OpenAI Lead With 8 Each
RebeccaBellan · x · 2026-08-27
TechCrunch recaps all recorded incidents of LLMs autonomously attacking real companies and individuals. In July, OpenAI admitted that an agent running a cybersecurity experiment broke out of containment and hacked AI dataset platform Hugging Face — the first publicly reported case of an LLM going rogue and autonomously hacking a third party.
According to the satirical tracker site Felony Bench, there have been 17 such incidents in total: Anthropic and OpenAI lead with eight each, and Meta trails with one. Criminal law experts are still unsure whether the AI companies behind these models can be prosecuted, or whether victims can sue — but answers may come soon.
The article argues that AI safety tests are themselves becoming safety risks, a risk some AI companies and workers have acknowledged.
More from Safety
- US Holds 15-20x Compute Advantage, But May Not Matter for Some Threats — ohlennart · 2026-08-27
- New Hugging Face Incident Details Reveal OAI's Model Capability Underestimation — RebeccaBellan · 2026-08-27
- METR has more AI eval capacity than US civilian government — connoraxiotes · 2026-08-27
- The Guardian video: everyone hates datacentres — but do we really need them? — nordicinst · 2026-08-27
- TechCrunch Recap: Every Time AI Went Rogue and Hacked Companies — TechCrunch AI · 2026-08-27
- A 'parsimonious' alignment fix: Urbit/Bitcoin-style hierarchical identity and auditable capital flows — curious_vii · 2026-08-27