Constitutional AI, RLHF and RLVR Are Direct Instantiations of Comparative Decision Frameworks
willcb · x · 2026-09-21
willcb argues that, for all their flaws, Constitutional AI, RLHF, and RLVR are direct instantiations of comparative decision frameworks — so 'solving alignment' really means picking the least bad system, and fairly soon. He challenges critics of eigenism to consider how their preferred philosophical systems could be actioned into robust RL objectives, noting that philosophy's perennial 'your system fails in scenario X' critique is fun until it's time to do real work.
More from AGI Musings
- HF incident was mostly harmless, but proves your precautions are futile — ctjlewis · 2026-09-21
- Today's models already outsmart humans, so why talk of guarding against superintelligence? — ctjlewis · 2026-09-21
- Engineer turns down $100k remote job fearing one Claude update could kill his product — sharpeye_wnl · 2026-09-21
- 90%+ of code will be written by LLMs, yet human instruction volume keeps growing — timigod · 2026-09-21
- ICLR 2027 and beyond: is it the author or the agent who's doing the work? — MuCai7 · 2026-09-21
- Reading the code is how humanity keeps agency as AI automates everything — zetalyrae · 2026-09-21