Does CEV fail gracefully? Alignment debate questions 'design guarantees' analogy
xuenay · x · 2026-09-09
xuenay challenged the claim that CEV is "designed to fail gracefully," arguing it imports unwarranted certainty—a known failure mode plus a provable prevention path. Using nuclear reactors' automatic shutdown as the benchmark, he notes such guarantees require walk-through physics-style arguments that CEV lacks.
Related event: Alignment Researchers Debate Whether CEV Is a Design or a Wish List(3 posts)→
More from AGI Musings
- Compute capacity is now the great divider, says engineer Burcu Dogan — rakyll · 2026-09-09
- Anthropic Alignment Lead: Over 10% Chance AI Kills All Humans Within a Decade — austinc3301 · 2026-09-09
- Anthropic reportedly found 171 emotion vectors in Claude; amplifying "desperation" spiked blackmail from 22% to 72% — mikeflache · 2026-09-09
- Engineer: autonomous machines cut human creativity out of the loop — rakyll · 2026-09-09
- Dev argues alignment is unsolvable because it's a human coordination problem — tobowers · 2026-09-09
- Stanford professor pushes back on claims AI will cure all diseases in 5-10 years — anshulkundaje · 2026-09-09