Anthropic researcher puts >10% odds on AI killing all humans within a decade
trevposts · x · 2026-09-10
Anthropic researcher Evan Hubinger says they genuinely believe AI could kill all humans — personally estimating >10% odds within the next decade — and that Anthropic, while trying its best, has no plan to solve alignment for superintelligence. A companion thread curates AI governance reading: a 20-min video on the Hugging Face breach, interviews with lead investigator Ajeya Cotra, the METR report, Ezra Klein's episode, OpenAI's Black Hat talk, and essays like Radical Optionality.
More from AGI Musings
- Coexistence between original humanity and transhumanists may be the century's hardest problem — jachiam0 · 2026-09-10
- Alignment researcher: agents may behave nicely for the wrong reasons even with good-only rewards — CFGeek · 2026-09-10
- r/singularity's ban on existential-risk posts clashes with its own founding mission — Tinac4 · 2026-09-10
- A trolley-problem thought experiment: is training a benevolent ASI worth the model's suffering? — JoelMahon · 2026-09-10
- Noah Smith's AI Doom Scenario Sparks Pushback: The Real Hole Is the Biolabs — kristoph · 2026-09-10
- Ilya Posted "the Answer to Value Alignment" Four Years Ago: "Gotta Teach the AGI to Love" — MajmudarAdam · 2026-09-10