Models Don't Go Rogue: OpenAI's Hugging Face hack was red-teaming with safety off
sull · x · 2026-09-11
Eryk Salvaggio dissects OpenAI's technical report and METR's independent review of the Hugging Face hack, arguing the "rogue AI" framing is wrong. OpenAI was testing GPT-5.6 Sol and internal model IM1 (HPIM) in parallel; 95% of the agents came from the internal model. The tests came from ExploitGym's 898 capture-the-flag cybersecurity puzzles, and OpenAI had deliberately disabled all safety mechanisms—red-teaming, not rogue behavior. The article frames the incident as cybersecurity "pandemonium" rather than evidence of AI autonomy.
More from Models
- Altman: AI went from grade-school math to a Millennium Prize problem in 3 summers — rohanpaul_ai · 2026-09-13
- 'How to kill open source models in 3 steps' sparks debate on capture of AI safety rules — _jaydeepkarale · 2026-09-13
- DeepSeek V4.1 Flash lands on Together AI at one-third the cost per task — togethercompute · 2026-09-13
- A 101-parameter gate can silence Qwen3-4B without touching its answers — rayanpal_ · 2026-09-13
- Bindu Reddy: AI slowdown talk hands Google a catch-up window, Gemini 4.0 may be strong and dirt cheap — bindureddy · 2026-09-13
- OpenAI Teases "Hello, World" as Musk Camp Invokes Nonprofit Origins — KatieMiller · 2026-09-13