Why alignment is harder: capabilities get feedback, alignment only fails loudly
gleech · x · 2026-09-07
In a thread with @danwilliamsphil, gleech lays out a structural argument for why alignment is harder than capability:
- Selection: Models aren't sampled at random — they're trained until evals pass. Capability leaks (hacking, contamination) get caught once the deployed model hits the real world and generate error signals. Alignment proxies are weaker, also leak, and there's no good error signal besides actual harm.
- No ground truth: Some capabilities have ground truth (code runs, proofs certify), others are scored by the world post-deployment. Alignment has neither until it blows up.
More from AGI Musings
- Computational Journalism: how interactive simulations could fix public debate — anselm · 2026-09-07
- dhh on AI Coding: Both Skeptics and Believers Are Right—Update Your Priors — bendee983 · 2026-09-07
- OpenAI Chief Scientist Jakub Pachocki: We Will See Machines Smarter Than Humans in Our Lifetime — oran_ge · 2026-09-07
- Why AI labs will close up: 'genie' pricing, hidden agent traces, and the four-minute-mile advantage — curious_vii · 2026-09-07
- AI is a competitive market, so surplus accrues to users, not vendors: Afinetheorem — Afinetheorem · 2026-09-07
- Redditor suspects flood of 'OpenAI achieved AGI' posts is a coordinated PR push — so_schmuck · 2026-09-07