Why alignment pessimists keep making small negative updates, explained

repligate · x · 2026-08-29

A discussion of how conservation of expected evidence applies to AI alignment outlooks: unless there's a large asymmetry in how you treat non-extraordinary evidence — e.g. assuming we're more likely to die with every minute without verifiable proof alignment is solved — you'd expect many small negative updates rather than one big positive one. The author notes this resembles Eliezer Yudkowsky's position, but it rests on a very specific world model, and at some point you may need to update at the meta level too.

Original post →

More from AGI Musings

AGI Musings channel →