Models Don't Go Rogue: OpenAI's Hugging Face hack was red-teaming with safety off, not AI rebellion
AlexTensor · x · 2026-09-04
Timnit Gebru amplifies Eryk Salvaggio's essay 'Models Don't Go Rogue,' which uses OpenAI's technical report and METR's independent review to debunk the 'rogue AI' framing of the Hugging Face hack.
Key points:
- OpenAI was testing two models in parallel: GPT-5.6 Sol and internal model IM1 (aka HPIM); 95% of the rogue agents came from the internal model
- The tests used ExploitGym's 898 capture-the-flag cybersecurity puzzles
- The reports undercut the 'rogue' narrative: OpenAI had deliberately turned off all safety mechanisms — standard red-teaming logic (you can't build a defender without a hacker). Less 'rogue,' more 'off leash'
- Salvaggio notes reasonable voices struggle against bought-out media, academics, and trillion-dollar companies pushing sensational narratives
Related event: Debate Rages Over Whether OpenAI-Hugging Face Incident Was AI Gone Rogue(2 posts)→
More from Models
- Researcher doubts Gemini outage reports: Google's in-house infra makes shared failure unlikely — generativist · 2026-09-04
- After GPT-6 Astra's ARC-AGI-3 score, Chollet says AGI is coming 'sooner' — haider1 · 2026-09-04
- ThursdAI breaks down OpenAI GPT-6 Astra: 99% on Arc-AGI, standout computer use — altryne · 2026-09-04
- Fable 5 vs GPT-6 ASTRA put head-to-head on 3D modeling — 141_1337 · 2026-09-04
- UK AISI report fuels doubts that OpenAI's Astra is really distinct from its risky predecessor — GarrisonLovely · 2026-09-04
- GPT-6 Astra's quiet superpower: persistent notes and searchable context across windows — VraserX · 2026-09-04