gleech argues alignment lags capabilities: capability errors surface, goal errors only show after harm
In an alignment discussion thread with philosopher Dan Williams, gleech offered a framework for why "capability is easier to learn and generalize than alignment," thereby explaining why AI alignment lags behind capability progress.
The core argument: models are not random samples—they are trained until they pass evaluations before release. Capability flaws (cheating, data contamination, etc.) get exposed once deployed into the real world, providing usable error signals; some capabilities even have ground truth (code runs, proofs verify). Alignment has neither—the proxy metrics for alignment are weaker and equally prone to leakage, and apart from "causing actual harm" there is no good error signal, and that signal is itself far too costly. So selection pressure acts on capability far more readily than on alignment.
Confirmed
- gleech highlights a key uncertainty: training-set behavior largely determines what a model "can do" (shortcuts and cheating aside), yet says almost nothing about what it "wants"—many goals are behaviorally indistinguishable; in the worst case, zero-shot deception is also possible
- He gives the coherent agent argument: agents correct their own false beliefs (because false beliefs harm their goals) but protect their goals from modification (because changing goals is usually bad for those goals), making value errors harder to correct than capability errors
- The fragility argument: small capability errors usually degrade gracefully, while goal errors can be arbitrarily bad
- Theoretical support: cites the ICML 2023 paper "Invariance in Policy Optimisation and Partial Identifiability in Reward Learning" (Skalse, Russell, et al.), showing reward functions are only partially identifiable, and learned goals change when the data source changes
- Empirical support: cites the ICML 2022 paper "Goal Misgeneralization in Deep Reinforcement Learning" (Langosco, Sharkey, Krueger, et al.), showing RL agents pursuing wrong goals while retaining their capabilities
Why it matters
- This discussion attributes "why alignment lags capability" to a fundamental difference between the two: capabilities have ground truth or real-world feedback, enabling effective selection pressure, while alignment lacks cheap error signals and its problems only surface after the fact. If the argument holds, simply scaling data and evaluations is unlikely to automatically solve alignment.
2026-09-07 ~ 2026-09-07 · 9 related posts
Primary sources
- Why capabilities beat alignment: selection pressure and generalization, per gleech — gleech · 2026-09-07
- gleech: capabilities have ground truth, alignment only shows after it blows up — gleech · 2026-09-07
- Why alignment is harder: capabilities get feedback, alignment only fails loudly — gleech · 2026-09-07
- [source] Alignment Has No Good Error Signal: Models Are Trained Until Evals Pass, Then Leaks Surface — gleech · 2026-09-07
- Underdetermination in Alignment: Training Data Pins Down Capability but Not Goals — gleech · 2026-09-07
- [source] Goal Misgeneralization in Deep RL: Agents Keep Their Skills but Pursue the Wrong Goal — gleech · 2026-09-07
- [source] ICML Paper: Reward Functions Are Only Partially Identifiable, Even With Infinite Data — gleech · 2026-09-07
- Why Capabilities May Be Easier Than Alignment: Agents Correct Beliefs but Protect Goals — gleech · 2026-09-07
- Capability errors degrade smoothly, goal errors can be arbitrarily bad — gleech · 2026-09-07