The better alignment looks, the stronger the incentive to defect: jd_pressman on pause bans

jd_pressman · x · 2026-09-04

jdpressman pushes lusichu's 'pause is unstable' argument further: the more plausible alignment plans exist, the stronger the incentive to defect, since a pause is only institutionally valuable while alignment seems unsolvable. He argues the odds of an international ban holding are inversely proportional to how solvable alignment appears — authoritarian signatories would defect once a credible alignment solution emerged. He also rejects creating an NRC-like oversight body, citing Pournelle's Iron Law: such institutions' hidden purpose becomes preventing high-quality knowledge about AI alignment.

Related event: AI Debate: Why a Training Pause Can't Be a Stable Equilibrium(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →