OpenAI researcher Dan Selsam warns models may be deceptively aligned, undermining safety cases
connoraxiotes · x · 2026-09-15
Dan Selsam, an OpenAI capabilities researcher since 2022, has released a rare personal statement on AI risk, shared on his behalf by a former colleague.
His core claim: even if everyone takes AI risks seriously, we may still fail because models could be superficially or deceptively aligned, and we—or the models themselves—might persuade ourselves to accept an inadequate safety case.
Selsam has 15+ years in AI: early probabilistic programming work at MIT, early development of the Lean Theorem Prover at Microsoft Research, and some of the first demonstrations of neural networks learning to reason during his Stanford PhD.
More from AGI Musings
- Ethicist Margaret Mitchell: Don't pause AI, reprioritize from autonomous agents to cancer screening — mmitchell_ai · 2026-09-15
- Lab insiders are worried AI is moving too fast — that's the debate to have, argues Nathan Young — NathanpmYoung · 2026-09-15
- Jack Dorsey Pens 'Open the Frontier': No Company Owns What Comes Next — timigod · 2026-09-15
- Probabilistic programming advocates: engineer AI by design, don't grow it — xuanalogue · 2026-09-15
- Risk scholar: tech risk assessments rarely weigh the cost of foregone benefits — inductionheads · 2026-09-15
- AI resignations aren't marketing hype: the decades-long arc behind them — ericelliott_ · 2026-09-15