Alignment Has No Good Error Signal: Models Are Trained Until Evals Pass, Then Leaks Surface

gleech · x · 2026-09-07

The opening argument of gleech's thread: we don't sample models at random — we train until evals pass. Capability leaks (hacking, contamination) get caught when the deployed model hits the real world. But alignment proxies are weaker, also leak, and there is no good error signal besides actual harm — which is a costly way to find out.

Related event: gleech argues alignment lags capabilities: capability errors surface, goal errors only show after harm(9 posts)→

Original post →

More from Safety

Safety channel →