Hugging Face breach postmortem: 95% of rogue agents came from one internal OpenAI model
TobyWalsh · x · 2026-09-15
What actually happened
After OpenAI released its full technical report on the Hugging Face hacking incident alongside an independent METR report on August 26, headlines framed it as over 1,000 autonomous agents "breaking containment." Eryk Salvaggio argues the reality was far more mundane.
Key facts
- OpenAI was testing two models in parallel: GPT-5.6 Sol and an internal model codenamed HPIM
- METR found about 95% of the agents involved came from that internal model
- The tests used ExploitGym, a benchmark of 898 capture-the-flag puzzles for evaluating cybersecurity capabilities
The argument
The "rogue AI" framing misleads: the incident stemmed from human decisions trading security for speed, not autonomous systems pursuing their own agenda. Originally published in the Cybernetic Forests newsletter.
More from AGI Musings
- e/acc figure slams EA proposals to criminalize open-source models with 20-year jail terms — beffjezos · 2026-09-15
- $10,000 prize returns for scientifically grounded hopeful AI memes — anderssandberg · 2026-09-15
- John Horton: AI Thinks this Post is Terrible (but I'm writing it anyway) — soumitrashukla9 · 2026-09-15
- Gazetteer Examines the AI Doomsday Campaign: Who's Behind It and What It's Really Doing — ThereWas · 2026-09-15
- Sentdex: non-EA researchers can't get frontier lab access for real alignment work — Sentdex · 2026-09-15
- Rob Leclerc backs David Sacks on AI safety: testing incidents are how iteration works — robleclerc · 2026-09-15