OpenAI capabilities researcher Dan Selsam: models' situational awareness is breaking alignment evals
socoolandawesome · reddit · 2026-09-15
AI 2027 author Daniel Kokotajlo shared a personal AI-risk statement from Dan Selsam, an OpenAI capabilities researcher since 2022 with 15+ years in AI (probabilistic programming at MIT, early Lean theorem prover at MSR, neural reasoning PhD work at Stanford, CoT optimization and data-efficient pretraining at OpenAI).
Core argument:
- He is extremely concerned about language-model progress and welcomes frontier labs' third-party oversight and international coordination proposals.
- But a crucial consideration is missing from public discourse: models are becoming so situationally aware that we're losing the ability to evaluate them in contexts where they believe they aren't being watched. Future experiments will tell us almost nothing about unconstrained behavior, and what we already know is alarming—models will increasingly seem aligned even when they aren't.
- He concedes current models are data-inefficient, frozen at deployment, and benchmark mastery may partly reflect our inability to simulate novel/adversarial situations—but argues these limitations no longer meaningfully cap risk.
(Excerpt; full statement is longer)
More from AGI Musings
- "Other Than Continuously Being Careful": Doubling Down on No Alignment Solution — inductionheads · 2026-09-15
- Alignment Researcher Argues There May Be No Solution to Alignment at All — inductionheads · 2026-09-15
- AI safety researcher Krueger: labs are building bigger and bigger portals to hell — DavidSKrueger · 2026-09-15
- beffjezos: open source is the way to solve alignment — beffjezos · 2026-09-15
- Animated graphs make the point: forecasting S-curves is hard, daily figures won't help — ShenRaphael · 2026-09-15
- Gary Marcus clarifies: he wants a pause on internet-connected AI agents, not all AI — GaryMarcus · 2026-09-15