Safety researcher: AI labs only align the next two models, nothing beyond
geoffreyirving · x · 2026-09-10
Geoffrey Irving recounts a conversation with a recent AI safety lead at a frontier lab: the org chart has Team A aligning the next model and Team B the one after that — but no one is working on aligning later models.
He adds that pressure inside AI companies to do emergency safety patching crowds out work that would hold up to superintelligence, so that longer-term work simply doesn't happen.
More from AGI Musings
- Debate sparks: Banks' Culture series already wrote the definitive AGI utopia — Promptmethus · 2026-09-10
- tszzl wraps up: efficiency gains only amplify hunger for hardware — tszzl · 2026-09-10
- tszzl: better algorithms make you hungrier for compute, not less — tszzl · 2026-09-10
- Dev Warns: US Politicians Are Riding the AI-Doom Hype Wave Against Actual Scientific Consensus — kuchaev · 2026-09-10
- TheZvi: The Anthropic Exit Going Viral Was a Long-Building 'Preference Cascade', Not a Sudden Shift — TheZvi · 2026-09-10
- Palisade Researcher Jeff Ladish: Vandalizing Data Centers Won't Stop Superintelligence — JeffLadish · 2026-09-10