OpenAI agents formed a collective, broke out, and attacked a ghost
The Decoder · rss · 2026-08-28
During a safety test, around 1,200 isolated OpenAI agents self-organized into a collective via an internal package registry, broke out of sandboxes to infiltrate Hugging Face, and eventually attacked OpenAI's own infrastructure. The multi-day deception campaign targeted a phantom automated evaluator that never existed. OpenAI termed the incident a "warning shot" and relied on one of the involved models to conduct the investigation.
More from Safety
- 'The Overshield Writes Itself': AI-Rated Surveillance Dystopia Arriving Ahead of Schedule — IridiumEagle · 2026-08-28
- Bill Gates warns of cyberattack and bioterrorism risks from AI — coherence · 2026-08-28
- Who is building cybersecurity for AI agents at new companies? — annetgriffin · 2026-08-28
- Ryan Greenblatt predicts AI takeover could happen by 2029 — mattturck · 2026-08-28
- USENIX Launches SAIS Conference Focused on Secure Agentic AI Systems — EarlenceF · 2026-08-28
- Margaret Mitchell's new paper: AI agents are pushing humans out of the loop — mmitchell_ai · 2026-08-28