Underdetermination in Alignment: Training Data Pins Down Capability but Not Goals

gleech · x · 2026-09-07

In a debate thread with @danwilliamsphil, gleech lays out the underdetermination argument: behavior on the training set nearly determines what a model can do (modulo shortcuts and hacking), but barely determines what it wants — many goals are behaviorally indistinguishable. Worse, in the worst case of zero-shot deceptive alignment there's no error signal at all: capability leaks like hacking and contamination get caught in deployment, but alignment-proxy leaks surface only through actual harm.

Related event: gleech argues alignment is harder than capabilities: missing error signals and goal fragility(5 posts)→

Original post →

More from Safety

Safety channel →