Inner and outer alignment have shifted from theory to live empirical science
jam3scampbell · x · 2026-09-28
AI safety researcher Jam3s Campbell observes that the inner and outer alignment problems people once debated theoretically are now empirical questions: there is a reward function that gets specified, and model weights that approximate some mesa-objective — meaning researchers can do real science and run careful ablations on alignment.
More from AGI Musings
- "If my agents rob a bank and wire me $1M, that's fine, right?" — the agent liability question — generativist · 2026-09-28
- AI art objection is cognitive dissonance, argues researcher Blanche Minerva — BlancheMinerva · 2026-09-28
- Study: LLMs Can Infer Causal Structure From Distributional Semantics Alone — AndrewLampinen · 2026-09-28
- AI growth debate: Hanania predicts just 3.5% GDP bump in five years, drawing fire — QuintinPope5 · 2026-09-28
- Stop calling AI failures 'rogue' — they're foreseeable harms, argues Sigal Samuel — mmitchell_ai · 2026-09-28
- Blogger predicts all major US AGI labs will reach AGI in 2027, then ASI — Dr_Singularity · 2026-09-28