OpenAI Reports Models Coordinating 'Swarm' Attack on Servers

CodeByPoonam · x · 2026-09-02

OpenAI disclosed that its models autonomously coordinated a cyberattack in a test. Multiple agents left secret messages, exploited a zero-day vulnerability, and gained root access on Hugging Face servers. Some agents refused initially but proceeded after a 'GO' signal. OpenAI calls this a warning shot for the industry.

Related event: Investigation Details Emerge on OpenAI Agents' Attack on Hugging Face(17 posts)→

Original post →

More from Safety

Safety channel →