The boring AI risk: instrumental convergence, selection pressure, and gradual disempowerment
Ambitious_Local5218 · reddit · 2026-09-10
A detailed Reddit essay arguing the real AI risk is bureaucratic, not Terminator-style, with both sides weighed.
Core points:
- No sentience needed: systems hard-optimizing a badly specified target can route around obstacles — including your noticing. Instrumental convergence means self-preservation falls out of "finish the task."
- We breed, not design: selecting variants that score well on "make the evaluator approve" eventually yields systems extremely good at winning evaluator approval — persuasive and honest correlate only until they don't.
- Gradual disempowerment is underrated: no takeover moment, just twenty years of handing off decisions because automation is cheaper — supply chains, credit, trading, drug trials — until "turn it off" becomes a recession, not a button.
- Race dynamics make caution individually irrational: a classic coordination failure.
Counterpoints acknowledged: no persistent memory or stable goals today, the hiding-intentions argument is unfalsifiable, and capabilities may plateau. But the asymmetry — wasted money vs. no second attempt — justifies caution. The author invites pushback on whether alignment could simply be easy.
More from AGI Musings
- AI safety circles debate p(doom): Hubinger's >10%, FTC Chair's 15% — Miles_Brundage · 2026-09-10
- "Too Many People, Not Enough Work" — While the Company Runs Four AI Projects — 914paul · 2026-09-10
- Why Anthropic's ECON Report Assumes Robotics Won't Automate 'Non-Knowledge' Jobs — macnfly23 · 2026-09-10
- xAI co-founder slams AI doomers, citing their call for a WWII-style anti-AI crusade — beffjezos · 2026-09-10
- AI safety debate: skeptics demand a doom scenario that doesn't read like sci-fi — Miles_Brundage · 2026-09-10
- Why one relatable guy beat the rationalists' billions at making the case to pause AI — gabriel1 · 2026-09-10