Constitutional AI, RLHF and RLVR Are Direct Instantiations of Comparative Decision Frameworks

willcb · x · 2026-09-21

willcb argues that, for all their flaws, Constitutional AI, RLHF, and RLVR are direct instantiations of comparative decision frameworks — so 'solving alignment' really means picking the least bad system, and fairly soon. He challenges critics of eigenism to consider how their preferred philosophical systems could be actioned into robust RL objectives, noting that philosophy's perennial 'your system fails in scenario X' critique is fun until it's time to do real work.

Related event: Researchers debate turning ethical philosophies into RL alignment objectives(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →