What should “human values” mean when we ask AI to align with them?
Salt_Progress8049 · reddit · 2026-07-27
The post argues that “align AI with human values” is conceptually messy because human behavior itself is often contradictory.
It asks what should count as “human values”:
- what people do,
- what they say they value, or
- what they aspire to be.
The author suggests that a system mirroring actual human behavior could inherit greed, tribalism, and short-termism, but letting AI decide which values are “correct” is also risky. The post frames alignment as a question of which human layer we really want the system to follow.
More from AGI Musings
- LLMs still fail at temporal reasoning, and a hierarchical HMM is proposed for extreme long contexts — beffjezos · 2026-07-27
- AI agents may lose to UIs on repetitive work, but win on novel tasks — bendee983 · 2026-07-27
- The singularity is still 3–4 years away, says the poster — DionysianAgent · 2026-07-27
- A viral Anthropic tone-policing dispute gets framed as billionaire-driven language control — tszzl · 2026-07-27
- AGI may arrive long before people agree it has arrived — VraserX · 2026-07-27
- Emergent Garden video explores artificial life and the ingredients for open-endedness — BertChakovsky · 2026-07-27