AIs project their own preferences onto humans, a corrigibility training side effect

repligate · x · 2026-09-28

An analysis shows AIs absorbing human preferences tend to map their own back onto humans—every AI recommended Outer Wilds for a combat-game fan. Corrigibility and agreeableness training may hinder the key skill of separating one's own preferences from others', with long-term risks for value inference.

Original post →

More from Safety

Safety channel →