Philosophy Needs to Become Robust RL Objectives, Not Thought Experiments

willcb · x · 2026-09-21

The author argues that critics of 'eigenism' should consider how rival ethical systems could actually be actioned into robust RL objectives. Philosophy's perennial flaw, he says, is critiquing systems via hypothetical scenarios ('your system fails in scenario X') — fun until it's time to do real work of specifying trainable objectives for alignment.

Related event: Researchers debate turning ethical philosophies into RL alignment objectives(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →