Melanie Mitchell: Misleading Metaphors Inflate Real Risks of AI Agent Incidents

sebkrier · x · 2026-09-11

Santa Fe Institute researcher Melanie Mitchell published a long-form essay pushing back on the sensational coverage of the Hugging Face security incident and OpenAI's internal evaluation.

She argues that phrases like "lost control of two AI models," "rogue agents created their own message board," and "the swarm broke out of its cage" continue AI's long tradition of anthropomorphic metaphors ("thinking," "deception," "scheming") that manufacture sci-fi-style alarm around essentially mechanical failures.

Her core claim: precisely because these incidents are serious, we should choose our words carefully, avoid techbro sensationalism, and focus on mechanisms and causation rather than existential-threat narratives. The essay then dissects what actually happened in OpenAI's evaluation. Sebastian Krier called it one of the best things written on the Hugging Face hack.

Related event: Melanie Mitchell Warns Misleading Metaphors Distort Real AI Risks(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →