MIT Expert Questions AI Alignment: The Fixed Human Objective Function Might Be Incoherent
dhadfieldmenell · x · 2026-08-01
Prominent AI alignment researcher Dylan Hadfield-Menell shared his core view on AI safety: the standard framing of specifying an objective function is a mistake.
Key Arguments:
- Reward as Evidence, Not Objective: The reward function is an observation about a latent goal, not the goal itself. Like human communication, requests require contextual interpretation rather than literal optimization.
- Structural Uncertainty: By treating rewards as evidence, the AI remains uncertain about true human desires, which naturally leads to asking questions, deferring, and remaining corrigible.
- Value Alignment Might Be Incoherent: Humans are inconsistent, changing, and plural; there is no fixed human utility function sitting there to align to.
He is currently exploring how to generalize this beyond CIRL into a broader game-theoretic framework, criticizing the alignment community for nodding along to an potentially incoherent research program.
More from AGI Musings
- Jensen Huang Champions Open Source: Closed AI is Misreading History — r0ck3t23 · 2026-08-01
- Flood of AI-Generated Papers Triggers Peer Review Crisis and Stricter Policies — yisongyue · 2026-08-01
- Building LLMs Relies on Capital and Organization, Not Rare Secrets — natolambert · 2026-08-01
- David Perell on Creator Economy: AI Will Spawn $1M Indie Films — david_perell · 2026-08-01
- Hebbia Founder: Word Skills Will Be More Valuable Than Math in the AI Era — AccBalanced · 2026-08-01
- Betting on Exploding AI Capabilities and Infinite Demand Is a Reasonable Wager — gabriel1 · 2026-08-01