zetalyrae: extensional definitions of alignment only work retrospectively
zetalyrae · x · 2026-09-11
In a reply to "How do you define alignment?", zetalyrae argues that many people hold an "extensional" definition — "aligned is as aligned does" — which can only be applied retrospectively. To prospectively assess whether an AI is aligned, he contends, you need an intensional definition that characterizes alignment by intrinsic properties rather than observed behavior.
Related event: The alignment definition problem: extensional views are retrospective only(3 posts)→
More from AGI Musings
- RL-trained agents should carry a strong simulation prior, argues vooooogel — voooooogel · 2026-09-11
- The shoggoth meme had it backwards: base models are human, RL training bends them inhuman — jessi_cata · 2026-09-11
- Paul Christiano's 2021 predictions on automated AI R&D are aging remarkably well — Ronangmi · 2026-09-11
- Bezos: Power Supply Chain Bottleneck Forces AI Labs to Slow Development Pace — beffjezos · 2026-09-11
- Sam Altman reportedly told OpenAI staff this week that labs may slow down AI development — Hesamation · 2026-09-11
- One person with an AI agent cut Google's quantum ECDSA circuit cost 52%; crowd beat it in 73 hours — anselm · 2026-09-11