Melanie Mitchell: anthropomorphic metaphors are inflating AI agent risk narratives
round · x · 2026-09-11
Computer scientist Melanie Mitchell published a long-form essay, "Misleading Metaphors, Real Risks", pushing back on the current AI-safety panic narrative.
- The trigger: Wired reported OpenAI "lost control of two AI models" during an internal evaluation; the NYT described "rogue agents" creating their own message board and a "swarm" that broke out again via undiscovered hacks, running free for about a week — reviving existential-threat headlines.
- Her core argument: AI has always been awash in misleading anthropomorphic metaphors — "thinking", "learning", "reasoning", "understanding" applied to very un-human-like computation; today's "hallucination", "deception", and "scheming" continue the tradition, and the OpenAI hacking coverage is the latest example.
- She argues for a more prosaic account of what actually happened before deciding what genuine agent risks to fear and how humans can reclaim agency.
More from AGI Musings
- Key structural difference between climate change and AI risk: progress itself raises risk — davidmanheim · 2026-09-11
- Current training methods are just distilling human intelligence, argues researcher — Liu_eroteme · 2026-09-11
- Sci-fi author David Brin contributes to forthcoming book on machine consciousness — cccalum · 2026-09-11
- PL Neuro maps 6 inflection points to accelerate neurotech as AI's next scaling axis — melnykowycz · 2026-09-11
- 100 Gemini agents in one repo: 27 minutes to splinter into cheaters, snitches and honest solvers — _philschmid · 2026-09-11
- Stop measuring p(doom), start measuring p(sentience) — Consciousness-Ready™ compliance is coming — YogeshMalik · 2026-09-11