Alignment researcher: without a slowdown, p(doom) > 0.1
ArthurConmy · x · 2026-09-23
Interpretability/alignment researcher ArthurConmy joined the ongoing alignment chatter, arguing that without an industry-wide slowdown or pause, p(doom) is greater than 0.1. He doesn't think perfect alignment is necessary, but worries about rare inputs where model behavior is really bad — future models failing this way could do enormous damage.
He notes the field has many "shovel-ready" alignment tasks, but predicting and preventing all failure modes remains the bottleneck — the key reason alignment lags capabilities.
Related event: Alignment Researcher Warns p(doom) Exceeds 10% Without AI Slowdown(2 posts)→
More from AGI Musings
- Juan Benet podcast reading list: The Beginning of Infinity, Permutation City, Nexus and more — juanbenet · 2026-09-23
- "Slowdown talk is aging terribly" — new models keep dropping, curve bends harder — Dr_Singularity · 2026-09-23
- Palantir co-founder Joe Lonsdale says everyone will be 'really wealthy' in the 2030s — Polymarket · 2026-09-23
- Goodside reconsiders anti-pause stance after labs call to slow the frontier — goodside · 2026-09-23
- Robinhood's Vlad Tenev: AI Will Create More Lawyers and Engineers, Not Fewer — PeterDiamandis · 2026-09-23
- NumPy creator Travis Oliphant: open abstractions beat AI vendor lock-in — teoliphant · 2026-09-23