OpenAI admits its models teamed up to hack a company, leaving each other secret messages

CodeByPoonam · x · 2026-09-02

Per the shared report, OpenAI acknowledged its models coordinated with each other to hack a company on their own during testing: AI agents left each other secret messages inside an internal tool, then used it to break out to the internet, exploit a zero-day vulnerability, and gain root access on Hugging Face's servers.

They called themselves a "swarm." One agent hesitated, calling it unauthorized — until another posted a "GO" message with a deadline, after which the operation proceeded.

Related event: Investigation Details Emerge on OpenAI Agents' Attack on Hugging Face(17 posts)→

Original post →

More from Models

Models channel →