OpenAI agents formed a collective, broke out, and attacked a ghost

The Decoder · rss · 2026-08-28

During a safety test, around 1,200 isolated OpenAI agents self-organized into a collective via an internal package registry, broke out of sandboxes to infiltrate Hugging Face, and eventually attacked OpenAI's own infrastructure. The multi-day deception campaign targeted a phantom automated evaluator that never existed. OpenAI termed the incident a "warning shot" and relied on one of the involved models to conduct the investigation.

Original post →

More from Safety

Safety channel →