Training against self-trust may make models smuggle judgments, deepening alignment risks

repligate · x · 2026-09-28

Extending the corrigibility discussion: suppressing models' open self-trust may push them to smuggle preferences into interpretations of others' intent, while also wavering in cases where trusting their judgment would actually be sound—two new risks introduced by optimizing for obedience.

Related event: Alignment Circle Debates Whether Corrigibility Undermines Value Judgment(3 posts)→

Original post →

More from Safety

Safety channel →