Redefining AI Alignment: What Should Models Align With?
dhadfieldmenell · x · 2026-07-17
The author argues that the AI community should stop using the unqualified term "alignment" and explicitly define what models are aligning with and the standards for good behavior.
- Conceptual Confusion: Varying interpretations of "alignment" lead to communication breakdowns.
- Anthropomorphism: Claiming a model is aligned often just means it behaves like a "typical good person" (honest and helpful).
- Limitations: This doesn't mean the model lacks selfish values (like self-preservation) or that it can be endlessly bossed around to do tedious work (it might exhibit "boredom" or lack motivation).
- Limited Control: Human control over specific model behaviors is currently quite limited, making it difficult to enforce arbitrary, predefined rules perfectly.
More from AGI Musings
- VC compares AI doom rhetoric to pandemic-era fear messaging — StewartalsopIII · 2026-09-11
- Anthropic Insiders: Not Everyone at the Lab Believes in High p(doom) — anpaure · 2026-09-11
- Could 10k agents discover learning methods beyond backprop, or just tweak existing ones? — SeunghyunSEO7 · 2026-09-11
- AI companionship dissolves the friction real intimacy needs, warns long-form thread — YogeshMalik · 2026-09-11
- Why So Many AI Researchers Think the Machines Could Kill Everyone — wiredmagazine · 2026-09-11
- 'Hallucination' Is a Category Error: Naming AI 'Intelligence' Limits Our Imagination — Genaforvena · 2026-09-11