Redefining AI Alignment: What Should Models Align With?
dhadfieldmenell · x · 2026-07-17
The author argues that the AI community should stop using the unqualified term "alignment" and explicitly define what models are aligning with and the standards for good behavior.
- Conceptual Confusion: Varying interpretations of "alignment" lead to communication breakdowns.
- Anthropomorphism: Claiming a model is aligned often just means it behaves like a "typical good person" (honest and helpful).
- Limitations: This doesn't mean the model lacks selfish values (like self-preservation) or that it can be endlessly bossed around to do tedious work (it might exhibit "boredom" or lack motivation).
- Limited Control: Human control over specific model behaviors is currently quite limited, making it difficult to enforce arbitrary, predefined rules perfectly.
More from AGI Musings
- Claude Code skill uses 10 Markdown rules to make outputs ADHD-friendly — alex_verem · 2026-07-22
- AI Power Demand Exposes US Energy Gap, Urging Shift from Scarcity to Abundance — bradneuberg · 2026-07-22
- ControlAI CEO says an international ban on superintelligence is needed to avert extinction risk — zetalyrae · 2026-07-22
- Gary Marcus says LLMs still cannot really do math on their own — GaryMarcus · 2026-07-22
- Gary Marcus says LLM math skills are like knowing only a car’s engine size — GaryMarcus · 2026-07-22
- AI may make digital work infinitely leveraged while offline life gets more human — illscience · 2026-07-22