Alignment is ill-defined, but monitorability and reversibility aren't
Afinetheorem · x · 2026-09-09
Drawing on an office conversation, the author argues that from a mechanism-design perspective "alignment" is hard to even define — a principal usually cannot specify a complete, stage-contingent optimum. By contrast, two related concepts are well-defined and operational: ex-post monitorability (can we observe and audit the system's behavior after the fact?) and reversibility ("pulling the plug" — can we stop or undo it when things go wrong?). The conclusion: alignment debates should put more weight on these concrete properties rather than chasing a formal definition of alignment.
More from AGI Musings
- Employees Who Ignore AI May Be the First to Be Laid Off — kevinsurace · 2026-09-09
- AI researchers debate tuning per-instance intelligence vs scaling population size — jd_pressman · 2026-09-09
- Terence Tao warns AI rumor-fueled crowds may end centuries of open science — GaryMarcus · 2026-09-09
- Are agent swarms more legible than single agents? Safety researchers debate — jd_pressman · 2026-09-09
- tszzl pushes back on agent-swarm doom: swarms may be more legible than single minds — tszzl · 2026-09-09
- Government AI strategist doubts models can fill a spreadsheet, researchers exasperated — AndyMasley · 2026-09-09