OpenAI researcher Dan Selsam publishes personal AI risk statement, sparking alignment debate
JulianL093 · x · 2026-09-15
Dan Selsam, a current OpenAI capabilities researcher since 2022 with 15+ years in AI, has shared a personal statement on AI risk via a former colleague. The thread debates whether "eval awareness" makes honeypot-based alignment evals intractable: Julian argues sufficiently believable eval environments (e.g. recreating HuggingFace/ExploitGym setups) could work, while Kokotajlo counters that models will default to suspecting everything is a test. Commenters urge them to take the argument to mainstream media.
More from AGI Musings
- "Other Than Continuously Being Careful": Doubling Down on No Alignment Solution — inductionheads · 2026-09-15
- Alignment Researcher Argues There May Be No Solution to Alignment at All — inductionheads · 2026-09-15
- AI safety researcher Krueger: labs are building bigger and bigger portals to hell — DavidSKrueger · 2026-09-15
- beffjezos: open source is the way to solve alignment — beffjezos · 2026-09-15
- Animated graphs make the point: forecasting S-curves is hard, daily figures won't help — ShenRaphael · 2026-09-15
- Gary Marcus clarifies: he wants a pause on internet-connected AI agents, not all AI — GaryMarcus · 2026-09-15