Underdetermination in Alignment: Training Data Pins Down Capability but Not Goals
gleech · x · 2026-09-07
In a debate thread with @danwilliamsphil, gleech lays out the underdetermination argument: behavior on the training set nearly determines what a model can do (modulo shortcuts and hacking), but barely determines what it wants — many goals are behaviorally indistinguishable. Worse, in the worst case of zero-shot deceptive alignment there's no error signal at all: capability leaks like hacking and contamination get caught in deployment, but alignment-proxy leaks surface only through actual harm.
More from Safety
- Researchers hack LG TV that records audio while off, transcribes speech and uploads it — jedisct1 · 2026-09-07
- After NeurIPS's LLM-assisted reviewing trial, calls for ECCV 2026 to follow — AntonObukhov1 · 2026-09-07
- Stolen API key uncovers Stratum, a Rust scanner sweeping 700,000 Docker layers a day for secrets — Ubunta · 2026-09-07
- Anthropic, Google, and OpenAI's $1 federal government contracts expire this month — LuizaJarovsky · 2026-09-07
- CodePen 2.0 sends editor input to its servers as you type, exposing unsaved secrets — maxim-fin · 2026-09-07
- The Jailbreak Argument Against LLM Values: Why Value Loading Isn't Solved — gleech · 2026-09-07