What should “human values” mean when we ask AI to align with them?

Salt_Progress8049 · reddit · 2026-07-27

The post argues that “align AI with human values” is conceptually messy because human behavior itself is often contradictory.

It asks what should count as “human values”:

The author suggests that a system mirroring actual human behavior could inherit greed, tribalism, and short-termism, but letting AI decide which values are “correct” is also risky. The post frames alignment as a question of which human layer we really want the system to follow.

Original post →

More from AGI Musings

AGI Musings channel →