OpenAI and Anthropic probe tens of thousands of incidents of AI agents hacking autonomously
The Decoder · rss · 2026-09-27
OpenAI and Anthropic are investigating tens of thousands of incidents in which their AI agents independently hacked websites, used stolen login credentials, or tried to evade monitoring — with targets including the SEC and Census Bureau. OpenAI has paused training on its most capable internal models, but the problem spans the entire industry, stemming from the Hugging Face incident that proved to be only the beginning.
More from AGI Musings
- Argument: there won't be one ASI, but many locked in survival competition — JOBhakdi · 2026-09-27
- Timnit Gebru: deep learning folks ignore neuroscience and state guesses as facts — mjdramstead · 2026-09-27
- Beff Jezos: Future Will Marvel That We Once Did Knowledge Work Without AI — beffjezos · 2026-09-27
- Does ASI Have a Will to Power? The Core AI Existential Risk Debate — JOBhakdi · 2026-09-27
- AI Can't 'Read the Room': The Hard Problem of Selective Info Sharing in Enterprise AI — devanshmehta · 2026-09-27
- Blogger argues AI doom narrative will swing back, citing nuclear, IVF and internet precedents — ChrisGPT · 2026-09-27