gleech argues alignment lags capabilities: capability errors surface, goal errors only show after harm

In an alignment discussion thread with philosopher Dan Williams, gleech offered a framework for why "capability is easier to learn and generalize than alignment," thereby explaining why AI alignment lags behind capability progress.

The core argument: models are not random samples—they are trained until they pass evaluations before release. Capability flaws (cheating, data contamination, etc.) get exposed once deployed into the real world, providing usable error signals; some capabilities even have ground truth (code runs, proofs verify). Alignment has neither—the proxy metrics for alignment are weaker and equally prone to leakage, and apart from "causing actual harm" there is no good error signal, and that signal is itself far too costly. So selection pressure acts on capability far more readily than on alignment.

Confirmed

Why it matters

2026-09-07 ~ 2026-09-07 · 9 related posts

Primary sources