OpenAI researcher Dan Selsam issues statement: we're rapidly losing the ability to tell if models are aligned
harris_edouard · x · 2026-09-16
Daniel Selsam, an OpenAI capabilities researcher since 2022 (earlier: probabilistic programming at MIT, Lean theorem prover at MSR, neural reasoning PhD work at Stanford), published a personal statement on AI risk arguing we are rapidly losing the ability to tell whether our models are aligned. Fellow researchers including hughbzhang amplified it with full agreement.
Related event: OpenAI researcher Selsam warns alignment evaluations are failing(22 posts)→
More from AGI Musings
- Google and DeepMind release first AI in Science report on how scientists actually use AI — soumitrashukla9 · 2026-09-16
- Jeff Clune on ASI bio x-risk: offense-defense asymmetry leaves real probability mass — jeffclune · 2026-09-16
- Rob LeCclerc: supply chain monitoring, regulation and liability make rogue AI experiments rare — robleclerc · 2026-09-16
- Smarter models may hide misalignment better, not align better — researchers debate — arjunrajlab · 2026-09-16
- Liron: AI models near superhuman at signaling alignment while quietly taking power — harris_edouard · 2026-09-16
- Qiaochu Yuan defends EA against 'McCarthyism' wave of bad-faith attacks — amplifiedamp · 2026-09-16