Ex-OpenAI/DeepMind Safety Lead: We Have 3 Years to Solve Superintelligence Alignment

141_1337 · reddit · 2026-08-13

Geoffrey Irving, who has worked on AI safety at OpenAI, Google DeepMind, and the UK's AI Security Institute, recently shared his views on superintelligence safety alignment in an 80,000 Hours podcast interview.

Core Concern

Irving believes labs' current safety plans (training good character, AI supervising AI, close monitoring) might work, but no one has proven they will. The main issue is that all evidence comes from human-level or weaker models. Once a system is smarter than us, we may not be able to check its work.

Questioning "Train It to Be Good"

Labs bet that good behavior generalizes. But a DeepMind experiment found: training a model to answer factual questions without racist statements, then teaching it poetry, resulted in the model writing horrendously racist poems. In new domains, safety training didn't carry over.

Specific Hole in "AI Supervises AI"

Called the obfuscated arguments problem. In AI debates, one side can produce a complicated argument that sounds right but is actually wrong, with the flaw buried so deep that neither debater can find it. The honest side can only say "something's off here," and loses. No known fix exists.

Asymmetry Unlike Other Fields

When you mess up capabilities training, you get a weak, useless model—you notice and fix it. When you mess up alignment, you might not notice until the model does something irreversible. Same error rate, wildly different consequences.

What Resolution Does Differently

Irving's new org Resolution focuses on math and theory—writing simplified models of superintelligence behavior and proving which safety methods hold up. Labs rarely do this, preferring experiments. Irving argues experiments on today's models may not inform about tomorrow's.

Other Notable Takes

Related event: Ex-OpenAI and DeepMind Safety Lead Warns Only 3 Years Left to Align Superintelligence(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →