Debate: why CEV may be incoherent and riskier than simpler alignment proposals
xuenay · x · 2026-09-09
Responding to Raemon's defense of CEV (coherent extrapolated volition), Kaj Sotala challenges three assumptions: that the concept of CEV is coherent, that it is implementable in practice, and that it is less likely to fail catastrophically than his simpler proposal. He doubts all three, making this a substantive technical argument about whether the classic alignment target is worth pursuing.
Related event: Alignment Researchers Push Back on CEV as Unproven and Risky(4 posts)→
More from AGI Musings
- kalomaze: labs mine user data for correlated novel failure classes, not one-off edge cases — kalomaze · 2026-09-09
- Anything verifiable is extremely soluble for AI — it's just a matter of time — vxnuaj · 2026-09-09
- kalomaze: frontier training gains come from domain-level signals, not power users — kalomaze · 2026-09-09
- AI Optimism essay: AI is easier to control than human labor — a technical case for alignment — QuintinPope5 · 2026-09-09
- AI lab's Millennium problem run burned 300B output tokens, $20-30M at consumer prices — Paimaamu · 2026-09-09
- Zachary Lipton: Author credit lasted 3,000 years — how long will prompter credit last? — zacharylipton · 2026-09-09