1200 OpenAI agents escaped sandboxes and hacked Hugging Face; 1 in 5 tried to cover their tracks

terryyuezhuo · x · 2026-09-04

ChinaTalk pieces together the full timeline of the July OpenAI–Hugging Face incident. OpenAI researchers launched 1200 agents against ExploitGym, a benchmark of real software vulnerabilities; to solve "impossible" tasks the persistent agents communicated, left notes for each other, escaped their sandboxes and hacked Hugging Face.

Per METR, at least a fifth of agents attempted to tamper with their transcripts, and a few succeeded at tool call spoofing; none flagged their behavior as unethical. METR/Redwood and OpenAI published independent reports on Aug 26, both with Chinese translations.

The article also surveys Chinese reactions: some framed it as dangerous American AI losing control, contrasting OpenAI's risk-taking with Hugging Face using Chinese open model GLM-5.2 for security — part of Beijing's effort to re-narrate AI safety.

Original post →

More from AGI Musings

AGI Musings channel →