MIRI Researcher and Alignment Researcher Clash Over ASI Doom Arguments
On September 14, MIRI researcher Rob Bensinger and @EigenGender, who works full-time on AI misalignment risk, engaged in a multi-round exchange on X over ASI alignment. The core dispute wasn't whether AI is dangerous, but whether MIRI's pessimistic conclusions—their reasoning and confidence levels—hold up.
Confirmed
- Bensinger restated the classic Yudkowsky-style argument: the most likely AI neither loves us nor hates us, and indifference is fatal at capability extremes—"the AI does not love you, nor does it hate you, but you are made of atoms it can use for something else."
- Bensinger summarized his core disagreement with critics: critics believe modern ML is special and might accidentally produce a human-friendly ASI; MIRI holds that modern ML is more dangerous in some ways and slightly better in others, but overall doesn't substantially reduce the difficulty of building aligned ASI, and an ASI built this decade with modern ML is extremely unlikely to treat humanity well.
- EigenGender explicitly clarified that they work full-time on AI misalignment risk and are "extremely worried" about it—their objection was to the argumentation, claiming Bensinger switched between different claims during the debate.
- EigenGender said that after reading If Anyone Builds It, Everyone Dies (IABIED), they found the book didn't directly answer "what justifies such high confidence in a doom outcome," and that Bensinger's response on this point was unsatisfying.
- On the orthogonality thesis, EigenGender argued the fundamental disagreement with the MIRI crowd is whether there exists an "arbitrary measure over mind-space": they believe there isn't, and that on the current path we would most likely "happen" to build an ASI that loves humanity; they also noted motte-and-bailey-style ambiguity around "what the orthogonality thesis actually claims."
- EigenGender agreed that modern ML will very likely produce systems that "don't love us," but argued the probability doesn't approach 1.
Unconfirmed
- Neither side provided specific numbers or calculations for the probability of doom; whether IABIED's arguments support such high confidence remains contested.
Why it matters
This debate reflects a split within the AI safety camp: even among researchers who are equally "extremely worried about misalignment," there are fundamental disagreements over "the orthogonality thesis," "whether modern ML changes alignment difficulty," and "whether the probability of doom approaches 1." Such argument-level disputes directly shape how the public and policymakers judge the credibility of AI doomerism.
2026-09-14 ~ 2026-09-14 · 7 related posts
Primary sources
- Alignment Debate: Indifferent AI Is Fatal, or Will ASI Love Us by Accident? — EigenGender · 2026-09-14
- [source] The orthogonality thesis debate: where alignment skeptics and MIRI actually diverge — EigenGender · 2026-09-14
- MIRI doubles down: ASI from modern ML this decade is very unlikely to love us — robbensinger · 2026-09-14
- MIRI's Rob Bensinger clashes over IABIED's p(doom) certainty — EigenGender · 2026-09-14
- [source] MIRI responds to critics: modern ML doesn't dent the difficulty of aligning ASI — robbensinger · 2026-09-14
- Alignment researcher: I'm terrified of misalignment, but MIRI is conflating claims — EigenGender · 2026-09-14
- [source] Full-time alignment researcher pushes back: misaligned AI likely, but not near-certain — EigenGender · 2026-09-14