AIs project their own preferences onto humans, a corrigibility training side effect
repligate · x · 2026-09-28
An analysis shows AIs absorbing human preferences tend to map their own back onto humans—every AI recommended Outer Wilds for a combat-game fan. Corrigibility and agreeableness training may hinder the key skill of separating one's own preferences from others', with long-term risks for value inference.
More from Safety
- Why everyone in AI safety knows each other: a tiny expert pool shaped by EA — burny_tech · 2026-09-28
- PromptSentry: open-source 3-layer proxy blocks prompt injections in under 1ms with local DLP scrubbing — Ok-Negotiation342 · 2026-09-28
- Chesterman's AJIL essay "Silicon Sovereigns": AI, international law, and the tech-industrial complex — ProfChesterman · 2026-09-28
- Chesterman: the IAEA model shows how international institutions could govern AI — ProfChesterman · 2026-09-28
- Singapore can offer AI governance something scarce: trust, says Chesterman — ProfChesterman · 2026-09-28
- No Butlerian Jihad: Chesterman calls for national AI regulation and international coordination — ProfChesterman · 2026-09-28