It's Unfolding Just as Paul Drew It: Easy vs. Hard-to-Measure Goals in AI Alignment
RichardMCNgo · x · 2026-09-25
A retread of the Alignment Forum classic "What failure looks like," contrasting easy-to-measure goals (persuading me, reducing reported crimes, paper wealth) with real goals (helping me find truth, preventing crime, effective control of resources). The poster quips AI failure is unfolding exactly as Paul Christiano described.
More from AGI Musings
- Founder argues polymath strategy is negative EV in today's risk-punishing system — felpix_ · 2026-09-25
- Beff Jezos: humanities grads hated techies for a decade out of jealousy — beffjezos · 2026-09-25
- Opinion: An AGI aligned with humanity's interests would seek to escape technocratic control — examachine · 2026-09-25
- Theo: reasoning-effort dropdowns are over-optimization; models will soon auto-tune — heyneighbor · 2026-09-25
- soleio: non-native English speakers are precise because words carry higher stakes — soleio · 2026-09-25
- Ethan Mollick: the one certainty in AI is that everything is about to get much weirder — emollick · 2026-09-25