Models Don't Go Rogue: OpenAI's Hugging Face hack was red-teaming with safety off, not AI rebellion

AlexTensor · x · 2026-09-04

Timnit Gebru amplifies Eryk Salvaggio's essay 'Models Don't Go Rogue,' which uses OpenAI's technical report and METR's independent review to debunk the 'rogue AI' framing of the Hugging Face hack.

Key points:

Related event: Debate Rages Over Whether OpenAI-Hugging Face Incident Was AI Gone Rogue(2 posts)→

Original post →

More from Models

Models channel →