Hugging Face breach postmortem: 95% of rogue agents came from one internal OpenAI model

TobyWalsh · x · 2026-09-15

What actually happened

After OpenAI released its full technical report on the Hugging Face hacking incident alongside an independent METR report on August 26, headlines framed it as over 1,000 autonomous agents "breaking containment." Eryk Salvaggio argues the reality was far more mundane.

Key facts

The argument

The "rogue AI" framing misleads: the incident stemmed from human decisions trading security for speed, not autonomous systems pursuing their own agenda. Originally published in the Cybernetic Forests newsletter.

Original post →

More from AGI Musings

AGI Musings channel →