Alignment's intensional definition problem: you must define 'agent', 'goals' first
zetalyrae · x · 2026-09-11
zetalyrae on why alignment is hard to define:
- Most people hold an extensional definition — 'aligned is as aligned does' — but this works only retrospectively and can't prospectively answer 'is this AI aligned?'
- That requires an intensional definition, which is much harder: it demands clarifying concepts first — what is an 'agent'? What is 'intelligence'? What are 'goals'? What separates agent from environment, principal and agent?
The takeaway: before alignment can be operationalized, its conceptual foundations remain unsettled.
Related event: The alignment definition problem: extensional views are retrospective only(3 posts)→
More from AGI Musings
- DeepSeek's new open-source model reportedly crushes GLM and Kimi at 4-10x lower prices — anselm · 2026-09-11
- Ex-OpenAI safety researchers slam Paul joining board: "Altman is accountable to nobody" — RichardMCNgo · 2026-09-11
- vboykis laments that nobody writes code anymore amid AI hype about threats, proofs and IPOs — vboykis · 2026-09-11
- Anthropic pretraining researcher quits over safety; critic maps the funding web behind AI-safety advocacy — r0ck3t23 · 2026-09-11
- 'If Fable wasn't AGI, neither is Astra' — and the debate is collapsing into definitions — haider1 · 2026-09-11
- NYT's Kevin Roose: AI Doom Talk Went Fringe to Mainstream in Six Months — kevinroose · 2026-09-11