Safety researcher: AI labs only align the next two models, nothing beyond

geoffreyirving · x · 2026-09-10

Geoffrey Irving recounts a conversation with a recent AI safety lead at a frontier lab: the org chart has Team A aligning the next model and Team B the one after that — but no one is working on aligning later models.

He adds that pressure inside AI companies to do emergency safety patching crowds out work that would hold up to superintelligence, so that longer-term work simply doesn't happen.

Related event: OpenAI researcher Irving warns firefighting safety patches crowd out superintelligence alignment work(5 posts)→

Original post →

More from AGI Musings

AGI Musings channel →