Irving: emergency safety patching pressure crowds out superintelligence-proof alignment work
geoffreyirving · x · 2026-09-10
Geoffrey Irving argues that pressure at AI companies to do emergency safety patching is too strong, so work designed to hold up to superintelligence never happens. He adds that more time could fund rigorous alignment attempts at Paul's day-job work and orgs like Resolution, and that pauses and reversals would better incentivize developers to aim for higher confidence.
More from AGI Musings
- AI safety frontier shifting from neural nets to mechanistic swarm interpretability — Hidenori8Tanaka · 2026-09-10
- Researcher leaves Google DeepMind for METR, puts AI takeover odds above 20% — sjgadler · 2026-09-10
- Tech companies will resemble hedge funds: tiny teams with AI agent leverage — brucemacv · 2026-09-10
- "If I claimed a 10% chance of killing people, the FBI would arrest me": AI safety jab goes viral — DavidSKrueger · 2026-09-10
- Be Grigori Perelman: the viral thread contrasting his refusal machine with AI agents chasing prizes — MrIvanAM · 2026-09-10
- Why now is the best time ever to found an AI safety startup, from alignment to governance — luke_drago_ · 2026-09-10