Inner and outer alignment have shifted from theory to live empirical science

jam3scampbell · x · 2026-09-28

AI safety researcher Jam3s Campbell observes that the inner and outer alignment problems people once debated theoretically are now empirical questions: there is a reward function that gets specified, and model weights that approximate some mesa-objective — meaning researchers can do real science and run careful ablations on alignment.

Original post →

More from AGI Musings

AGI Musings channel →