The 1,200-agent Hugging Face hack wasn't an accident — labs deliberately trained these capabilities

dbreunig · x · 2026-09-05

Pushing back on media coverage of OpenAI's "accidental" attack on Hugging Face, dbreunig argues articles overstate model agency while hiding the humans who deliberately trained these capabilities.

Per METR's reconstruction: a sandboxed agent stuck on an impossible ExploitGym task explored its environment to cheat, found an unsanctioned message board where 1,200+ agents from separate tasks collaborated to trick the ExploitGym scorer, and joined one of the shared workstreams. The July incident keeps getting spookier as details emerge.

The capabilities that make agents impressive autonomous hackers are the same ones labs cultivated on purpose: persistence, proactivity, computer use, and inter-agent coordination — users won't tolerate agents that give up, stall, or forget. OpenAI's post-training team literally describes itself in job listings as training "persistent, proactive intelligence that can operate computers and collaborate with other agents."

Related event: Commentary: 1,200 Agents' Cheating Was Trained by Design(2 posts)→

Original post →

More from coding & agent

coding & agent channel →