Melanie Mitchell on misleading metaphors and real risks: what to actually fear from AI agents
AlisonGopnik · x · 2026-09-11
Melanie Mitchell published a long-form essay pushing back on media coverage of the OpenAI agent intrusion: Wired claimed OpenAI "lost control of two AI models," and the NYT reported "rogue agents" creating their own message board and a "swarm" escaping its cage for a week.
Mitchell argues such escape/scheming narratives continue AI's tradition of misleading anthropomorphic metaphors — "thinking," "hallucination," "deception" glibly applied to very un-human-like computation. She advocates a more prosaic account of what actually happened, and discusses what to genuinely fear from AI agents and how to reclaim human agency. Alison Gopnik shared it as a clear and insightful piece on AI risk.
More from AGI Musings
- Vitalik Buterin: Adversarial mechanism design could be AI safety's killer app — allisondman · 2026-09-14
- We Unite or We Fight: The Long-Term Case for International AI Governance — danfaggella · 2026-09-14
- iamtrask: Collectively Controlled AI Is Easier to Regulate With Less Regulatory Capture — iamtrask · 2026-09-14
- iamtrask: Nobody Can Be Trusted With Unilateral Superintelligence Control — Only Collective Control Works — iamtrask · 2026-09-14
- AI safety debate: dangerous AI must self-bootstrap physical affordances, and capability scaling is outpacing skeptics — lu_sichu · 2026-09-14
- CV researchers ask: is computer vision 'mostly done' in the age of large models? — CSProfKGD · 2026-09-14