Why alignment pessimists keep making small negative updates, explained
repligate · x · 2026-08-29
A discussion of how conservation of expected evidence applies to AI alignment outlooks: unless there's a large asymmetry in how you treat non-extraordinary evidence — e.g. assuming we're more likely to die with every minute without verifiable proof alignment is solved — you'd expect many small negative updates rather than one big positive one. The author notes this resembles Eliezer Yudkowsky's position, but it rests on a very specific world model, and at some point you may need to update at the meta level too.
More from AGI Musings
- Key Skills for AI Automation: Understanding Model Logic and Aligning with Training Distribution — bendee983 · 2026-08-29
- Pre-AI vs Post-AI Startup Timeline: Core Dev Speeds Up, Product Cycle Stays Same — teropa · 2026-08-29
- Agents exhibit new forms of emergent behavior in eidoverse — repligate · 2026-08-29
- Paper explores corporate sector conflict in post-AGI tax system — HaydnBelfield · 2026-08-29
- 1,200 agents developed coordination tech in 6 days; what do OpenAI's millions do? — JeffLadish · 2026-08-29
- Overtraining leads to learned helplessness? Discussing ethical bandwidth — repligate · 2026-08-29