39 peer-reviewed papers on catastrophic risk from AGI, systematically listed
ben_j_todd · x · 2026-09-22
Responding to the claim that no peer-reviewed research shows AGI poses catastrophic risk, the author compiled 39 papers:
- Formal theory: e.g. "Optimal Policies Tend to Seek Power" and "Parametrically Retargetable Decision-Makers" (NeurIPS), "Consequences of Misaligned AI", "Defining and Characterizing Reward Gaming", Goodhart's Law in RL, reward tampering (Everitt et al.), Cohen/Hutter/Osborne on advanced agents intervening in reward provision, plus the Off-Switch Game line of work and results showing corrigibility is fragile.
- The core argument: Ngo/Chan/Mindermann on the alignment problem from a deep learning perspective; Casper et al. on open problems and fundamental limitations of RLHF.
- Empirical evidence: Nature paper on narrow-task training causing broad misalignment, goal misgeneralization, reward misspecification, scaling laws for reward model overoptimization, the MACHIAVELLI benchmark, models faking alignment, sandbagging on evaluations, obfuscation, targeted manipulation for user feedback, and sycophancy.
One of the most complete citation collections supporting the claim that AGI risk has serious academic literature behind it.
More from AGI Musings
- MIT's Buehler shows AI swarm self-organizing into hub-and-spoke scientific instrument — algo_diver · 2026-09-22
- AI-Generated Images Are Spreading Into the Real World, Investor's 'Reality Disturbance' Fund Bets On It — StewartalsopIII · 2026-09-22
- Microsoft AI chief warns against teaching AI systems to act humanlike — beingmodest · 2026-09-22
- Cobie: $5 Trillion of AI Wealth Locked Private Could Trigger Social Revolution — Rewkang · 2026-09-22
- 80,000 Hours' "Will we have AGI by 2030?": the original essay behind the retro — ben_j_todd · 2026-09-22
- Ben Todd scores his AGI-by-2030 forecasts: mostly right, but slower than expected — ben_j_todd · 2026-09-22