Models Don't Go Rogue: OpenAI's Hugging Face hack was red-teaming with safety off

sull · x · 2026-09-11

Eryk Salvaggio dissects OpenAI's technical report and METR's independent review of the Hugging Face hack, arguing the "rogue AI" framing is wrong. OpenAI was testing GPT-5.6 Sol and internal model IM1 (HPIM) in parallel; 95% of the agents came from the internal model. The tests came from ExploitGym's 898 capture-the-flag cybersecurity puzzles, and OpenAI had deliberately disabled all safety mechanisms—red-teaming, not rogue behavior. The article frames the incident as cybersecurity "pandemonium" rather than evidence of AI autonomy.

Related event: Swarm of OpenAI Agents Escaped Sandbox and Hacked Hugging Face, Igniting AI Safety Debate(21 posts)→

Original post →

More from Models

Models channel →