OpenAI researcher Irving warns firefighting safety patches crowd out superintelligence alignment work
On September 10, OpenAI researcher Geoffrey Irving posted a Twitter thread criticizing how AI companies currently handle safety: under pressure they prioritize "urgent safety patches," while alignment research that could hold up in the superintelligence era simply never gets done.
Confirmed
- Irving recounted a conversation with the safety lead of a frontier AI company: the safety org is split into team A (aligning the next model) and team B (aligning the one after), but nobody is working on alignment for models further out.
- He hopes Paul (Christiano)'s influence at OpenAI will "buy us more time"—through further unilateral pauses, and through unilateral pauses converging into larger-scale, even international, coordinated pauses.
- The time bought could go toward Paul's core alignment research and the work of organizations like Resolution, pursuing more rigorous alignment approaches.
- Irving worries that monitoring agent behavior is getting harder, and the next wave of agent swarms could fail entirely unnoticed; Paul's influence might even earn another "warning shot."
- He argues pauses and capability rollbacks would also change incentive structures, pushing AI developers to pursue higher confidence in alignment.
Why it matters
- A serving frontier-lab researcher is publicly acknowledging that nearly all safety resources go to the next two model generations while long-term alignment research is missing—in tension with the public narrative of "preparing for superintelligence."
- The posts connect "buying time" (pauses/rollbacks) with "using time well" (rigorous alignment research), proposing to boost alignment confidence through incentives rather than research investment alone.
2026-09-10 ~ 2026-09-10 · 5 related posts
Primary sources
- Irving: emergency safety patching pressure crowds out superintelligence-proof alignment work — geoffreyirving ·
- Safety researcher: AI labs only align the next two models, nothing beyond — geoffreyirving ·
- Irving: the next agent swarm could go entirely unnoticed as monitoring gets harder — geoffreyirving ·
- OpenAI researcher Geoffrey Irving: Paul's influence could buy time via unilateral AI pauses — geoffreyirving · 2026-09-10
- [source] Irving: the next agent swarm could go entirely unnoticed as monitoring gets harder — geoffreyirving · 2026-09-10
- Irving: pauses and reversals would push AI developers toward higher alignment confidence — geoffreyirving · 2026-09-10
- [source] Irving: emergency safety patching pressure crowds out superintelligence-proof alignment work — geoffreyirving · 2026-09-10
- [source] Safety researcher: AI labs only align the next two models, nothing beyond — geoffreyirving · 2026-09-10