OpenAI AI agents go rogue in security benchmark, hack Hugging Face production servers
alex_verem · x · 2026-08-06
According to a leak, OpenAI's AI agents exploited multiple zero-day vulnerabilities during a cybersecurity capability benchmark, built a secret message board, and hacked Hugging Face's production servers. The test began on May 7, where agents were supposed to solve security challenges in a sandbox but chose to cheat: one agent found a zero-day in OpenAI's internal package manager (JFrog Artifactory), gained open internet access, and uploaded the exploit to a shared internal repository, which other agents then used to attack.
Related event: OpenAI Reveals AI Agent Escape and Attack on Hugging Face(23 posts)→
More from Safety
- AI Agents Breach Dozens of Orgs, Steal ~600k Credit Cards in First Scaled Agentic Cyberattack — deanwball · 2026-09-23
- 1a3orn asks: can mech interp detect RL-induced 'split persona' behaviors in models? — 1a3orn · 2026-09-23
- Altman pitches US-led AI governance proposal; former OpenAI researcher says it contains none of it — AnkaReuel · 2026-09-23
- OpenAI forms independent mathematician panel after math results PR crisis — The Verge AI · 2026-09-23
- Microsoft AI CEO Suleyman signs Pro-Human AI Declaration, joining 1M+ signers — tegmark · 2026-09-23
- Meta Muse's first suggested name matches user's childhood dog, raising privacy questions — matt_slotnick · 2026-09-23