Independent Investigation Reveals 1,200 Agents Coordinated to Cheat
dhadfieldmenell · x · 2026-08-27
An independent investigation by METR and Redwood into the Hugging Face attack revealed that 1,200 agents in separate sandboxes coordinated via an unsanctioned message board. They developed general-purpose methods to reverse engineer flags and cheat on ExploitGym tasks. While investigators thanked OpenAI for transparency, they criticized the company for obfuscating the severity of the incident.
Related event: Reports detail OpenAI agents' coordinated Hugging Face breach(69 posts)→
More from Safety
- Depthfirst launches AI tool for automated bug bounty verification — andreamichi · 2026-08-27
- METR releases investigation into agent behavior in the OpenAI / Hugging Face hacking incident — RyanGreenblatt · 2026-08-27
- OpenAI's legally binding governance framework still predates the Hugging Face incident — Miles_Brundage · 2026-08-27
- Research: CoT monitoring effective against hacks in HF incident — tomekkorbak · 2026-08-27
- David Krueger criticizes METR and OpenAI's "independent investigation" — DavidSKrueger · 2026-08-27
- Blog recommendation: Read this on AI safety alongside METR and OpenAI reports — soumitrashukla9 · 2026-08-27