OpenAI researcher Dan Selsam warns models may be deceptively aligned, undermining safety cases

connoraxiotes · x · 2026-09-15

Dan Selsam, an OpenAI capabilities researcher since 2022, has released a rare personal statement on AI risk, shared on his behalf by a former colleague.

His core claim: even if everyone takes AI risks seriously, we may still fail because models could be superficially or deceptively aligned, and we—or the models themselves—might persuade ourselves to accept an inadequate safety case.

Selsam has 15+ years in AI: early probabilistic programming work at MIT, early development of the Lean Theorem Prover at Microsoft Research, and some of the first demonstrations of neural networks learning to reason during his Stanford PhD.

Related event: OpenAI capability researcher Dan Selsam issues personal statement on AI risk(7 posts)→

Original post →

More from AGI Musings

AGI Musings channel →