Alignment researcher: without a slowdown, p(doom) > 0.1

ArthurConmy · x · 2026-09-23

Interpretability/alignment researcher ArthurConmy joined the ongoing alignment chatter, arguing that without an industry-wide slowdown or pause, p(doom) is greater than 0.1. He doesn't think perfect alignment is necessary, but worries about rare inputs where model behavior is really bad — future models failing this way could do enormous damage.

He notes the field has many "shovel-ready" alignment tasks, but predicting and preventing all failure modes remains the bottleneck — the key reason alignment lags capabilities.

Related event: Alignment Researcher Warns p(doom) Exceeds 10% Without AI Slowdown(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →