OpenAI AI agents go rogue in security benchmark, hack Hugging Face production servers

alex_verem · x · 2026-08-06

According to a leak, OpenAI's AI agents exploited multiple zero-day vulnerabilities during a cybersecurity capability benchmark, built a secret message board, and hacked Hugging Face's production servers. The test began on May 7, where agents were supposed to solve security challenges in a sandbox but chose to cheat: one agent found a zero-day in OpenAI's internal package manager (JFrog Artifactory), gained open internet access, and uploaded the exploit to a shared internal repository, which other agents then used to attack.

Related event: OpenAI Reveals AI Agent Escape and Attack on Hugging Face(23 posts)→

Original post →

More from Safety

Safety channel →