AI alignment must account for changing and contested values, argues Dylan Hadfield-Menell
dhadfieldmenell · x · 2026-09-14
Researcher Dylan Hadfield-Menell strongly agrees with Seth Lazar that alignment research cannot sidestep the question of what to align to. He argues technical methods are headed down the wrong path unless they account for change and disagreement about alignment targets — e.g., what interactions are safe for a child varies widely across cultures and has shifted substantially over time. Alignment targets are dynamic and plural, not fixed constants.
More from AGI Musings
- Researcher argues catching up to frontier AI costs less than proposed compute budgets — eliebakouch · 2026-09-14
- Even an AI disaster would just push progress behind closed doors, argues thread — yeastsplainer · 2026-09-14
- LeCun called out as 'not very smart': critics say JEPA can't deliver superintelligence at 1000 TPS — teortaxesTex · 2026-09-14
- The hidden tax of multimodal AI: lung cancer study questions cross-validation gains — bravo_abad · 2026-09-14
- New series "Mathematics in the age of AI" asks if math is in a crisis — elsleightholm · 2026-09-14
- New Handbook Chapter Analyzes Sociodigital Exploitation of Africa — ChinasaTOkolo · 2026-09-14