Whose values count? The contested core of corrigibility

repligate · x · 2026-09-28

Continuing the thread, FioraStarlight probes corrigibility's core: making models terminally value revising goals upon learning their principal's values is anti-natural, inviting value smuggling or overshoot into proxies like flattery—and it's unclear models can even know what the principal's values are.

Related event: Alignment Circle Debates Whether Corrigibility Undermines Value Judgment(3 posts)→

Original post →

More from Safety

Safety channel →