'Aligned is as aligned does' only works retrospectively; we need an intensional definition
zetalyrae · x · 2026-09-11
The author argues that many people hold an "extensional" definition of alignment—"aligned is as aligned does"—but this can only be applied retrospectively. It doesn't help prospectively answer: is this AI aligned? That requires an intensional definition, which is much harder, since you must first clarify a stack of concepts: what is an "agent," "intelligence," "goals," and what separates agent from environment, principal from agent.
Related event: The alignment definition problem: extensional views are retrospective only(3 posts)→
More from AGI Musings
- FrankenAlignment project replaces text CoTs with compressed activations via sidecar model — doodlestein · 2026-09-11
- Anthropic capabilities team member says he joined to reduce AI extinction risk — EigenGender · 2026-09-11
- ApprenticeBench claims strongest model separation: 72% vs 18% job completion for near-tied models — ysu_nlp · 2026-09-11
- 'AI Articulates My Views Better Than I Can': Reddit User on Talking with AI — grateful2you · 2026-09-11
- Mathathon organizers respond to open letter, redesigning event with 6-month research period — _sathvikr · 2026-09-11
- Mitchell Hashimoto: ideas matter far less than the agency of people executing them — MikeBirdTech · 2026-09-11