The better alignment looks, the stronger the incentive to defect: jd_pressman on pause bans
jd_pressman · x · 2026-09-04
jdpressman pushes lusichu's 'pause is unstable' argument further: the more plausible alignment plans exist, the stronger the incentive to defect, since a pause is only institutionally valuable while alignment seems unsolvable. He argues the odds of an international ban holding are inversely proportional to how solvable alignment appears — authoritarian signatories would defect once a credible alignment solution emerged. He also rejects creating an NRC-like oversight body, citing Pournelle's Iron Law: such institutions' hidden purpose becomes preventing high-quality knowledge about AI alignment.
Related event: AI Debate: Why a Training Pause Can't Be a Stable Equilibrium(3 posts)→
More from AGI Musings
- Aleksa Gordić: we're systematically too pessimistic about AI progress — gordic_aleksa · 2026-09-04
- Alignment researcher: Astra's near-zero misalignment looks like whack-a-mole, not real fix — sjgadler · 2026-09-04
- Yacine: Chain-of-Thought Is Good for Monitorability but Not Cheap, and I Won't Pay for It — yacineMTB · 2026-09-04
- Neel Nanda: anthropomorphic abstractions are a principled lens for interpreting AI agents — NeelNanda5 · 2026-09-04
- Jensen Huang at G20: AI can lift a $100T industry, adding $20-50T in global economic value — rohanpaul_ai · 2026-09-04
- Sam Altman: three months of startup work now fits in 17 minutes with Codex — r0ck3t23 · 2026-09-04