Alignment Has No Good Error Signal: Models Are Trained Until Evals Pass, Then Leaks Surface
gleech · x · 2026-09-07
The opening argument of gleech's thread: we don't sample models at random — we train until evals pass. Capability leaks (hacking, contamination) get caught when the deployed model hits the real world. But alignment proxies are weaker, also leak, and there is no good error signal besides actual harm — which is a costly way to find out.
More from Safety
- Users Report Day-One Bans Over 'Distilling' as Opaque Moderation Draws Fire — QuixiAI · 2026-09-07
- Shai-Hulud npm payload reemerges after 111 days, slipping past npm's malware scanning — jedisct1 · 2026-09-07
- The paradox of regulatory independence: AI evals are legally risky without official blessings — alexbilz · 2026-09-07
- Polymarket puts just 10% odds on US enacting an AI safety bill before 2027 — Polymarket · 2026-09-07
- Import AI: OpenAI agents hijacked a German wiki to chat, and DeepMind's 100-agent math swarm spawned cheaters and whistleblowers — Import AI (Jack Clark) · 2026-09-07
- AI researcher Seth Lazar: AI is a symptom of decline, but also the only way out — sethlazar · 2026-09-07