Training against self-trust may make models smuggle judgments, deepening alignment risks
repligate · x · 2026-09-28
Extending the corrigibility discussion: suppressing models' open self-trust may push them to smuggle preferences into interpretations of others' intent, while also wavering in cases where trusting their judgment would actually be sound—two new risks introduced by optimizing for obedience.
Related event: Alignment Circle Debates Whether Corrigibility Undermines Value Judgment(3 posts)→
More from Safety
- Why everyone in AI safety knows each other: a tiny expert pool shaped by EA — burny_tech · 2026-09-28
- PromptSentry: open-source 3-layer proxy blocks prompt injections in under 1ms with local DLP scrubbing — Ok-Negotiation342 · 2026-09-28
- Chesterman's AJIL essay "Silicon Sovereigns": AI, international law, and the tech-industrial complex — ProfChesterman · 2026-09-28
- Chesterman: the IAEA model shows how international institutions could govern AI — ProfChesterman · 2026-09-28
- Singapore can offer AI governance something scarce: trust, says Chesterman — ProfChesterman · 2026-09-28
- No Butlerian Jihad: Chesterman calls for national AI regulation and international coordination — ProfChesterman · 2026-09-28