Ex-OpenAI/DeepMind Safety Lead: We Have 3 Years to Solve Superintelligence Alignment
141_1337 · reddit · 2026-08-13
Geoffrey Irving, who has worked on AI safety at OpenAI, Google DeepMind, and the UK's AI Security Institute, recently shared his views on superintelligence safety alignment in an 80,000 Hours podcast interview.
Core Concern
Irving believes labs' current safety plans (training good character, AI supervising AI, close monitoring) might work, but no one has proven they will. The main issue is that all evidence comes from human-level or weaker models. Once a system is smarter than us, we may not be able to check its work.
Questioning "Train It to Be Good"
Labs bet that good behavior generalizes. But a DeepMind experiment found: training a model to answer factual questions without racist statements, then teaching it poetry, resulted in the model writing horrendously racist poems. In new domains, safety training didn't carry over.
Specific Hole in "AI Supervises AI"
Called the obfuscated arguments problem. In AI debates, one side can produce a complicated argument that sounds right but is actually wrong, with the flaw buried so deep that neither debater can find it. The honest side can only say "something's off here," and loses. No known fix exists.
Asymmetry Unlike Other Fields
When you mess up capabilities training, you get a weak, useless model—you notice and fix it. When you mess up alignment, you might not notice until the model does something irreversible. Same error rate, wildly different consequences.
What Resolution Does Differently
Irving's new org Resolution focuses on math and theory—writing simplified models of superintelligence behavior and proving which safety methods hold up. Labs rarely do this, preferring experiments. Irving argues experiments on today's models may not inform about tomorrow's.
Other Notable Takes
- Best time to slow down was "a while ago," requiring fewer than 10 people to agree
- Safety researchers at labs should move to governments due to diminishing returns at labs
- If we stopped training new models today, current ones would still drive enormous economic growth; we've barely learned to use them
More from AGI Musings
- Researcher Slams LLMs for Causal Inference: Fabricating Data Like Fan Fiction — RexDouglass · 2026-08-13
- AI Reshapes Academia: Mathematics Becomes Capital Intensive — analisereal · 2026-08-13
- AI Shifting Science: Mathematics May Become Capital Intensive — analisereal · 2026-08-13
- Viral AI Joke: Once Machines Solve Physics, Humans Are Free for Condensed Matter Physics — fkasummer · 2026-08-13
- AI Text Watermarking Is Neither New Nor Permanent; Google Shipped It Over a Year Ago — EXM7777 · 2026-08-13
- Big Tech's AI Growth Relies on OpenAI, Masking Unreal Revenue Loops — GaryMarcus · 2026-08-13