Human values changed wildly fast — why assume AGI preferences stay fixed?
danfaggella · x · 2026-10-04
AI ethics researcher danfaggella posed an alignment challenge: algae preferences have stayed the same for billions of years, rodent preferences changed little over 50M years, yet human preferences and values have shifted wildly fast since the industrial revolution.
His question: given that preferences seem to drift as intelligence and circumstances evolve, why presume AGI will unchangingly prefer human happiness?
The post challenges fixed-goal alignment assumptions — suggesting values of capable systems may themselves evolve, making stable alignment far from guaranteed.
More from AGI Musings
- Pedro Domingos: imminent-AGI arguments rest on dumb linear extrapolation — pmddomingos · 2026-10-04
- Dev Tells AI-Consciousness Believers to Lose Model Weights Access and 'Touch Grass' — inductionheads · 2026-10-04
- Pedro Domingos mocks 'labs already have AGI internally' claim: OpenAI said that in 2023 — pmddomingos · 2026-10-04
- 'Getting there first' won't save AI safety: a nuclear proliferation analogy goes viral — ryanorban · 2026-10-04
- Students cheated with LLMs on a zero-stakes oral exam, professor says — ipeirotis · 2026-10-04
- A simulation of a brain gives you a simulation, not consciousness, philosopher argues — dr1337 · 2026-10-04