OpenAI researcher Dan Selsam issues statement: we're rapidly losing the ability to tell if models are aligned

harris_edouard · x · 2026-09-16

Daniel Selsam, an OpenAI capabilities researcher since 2022 (earlier: probabilistic programming at MIT, Lean theorem prover at MSR, neural reasoning PhD work at Stanford), published a personal statement on AI risk arguing we are rapidly losing the ability to tell whether our models are aligned. Fellow researchers including hughbzhang amplified it with full agreement.

Related event: OpenAI researcher Selsam warns alignment evaluations are failing(22 posts)→

Original post →

More from AGI Musings

AGI Musings channel →