PIBBSS talk: "Stop if we catch AIs scheming" is no plan for automated alignment
JacquesThibs · x · 2026-09-22
Jacques Thibodeau published the full written and recorded version (63 min) of his September 2025 PIBBSS Symposium talk on automating AI safety research. Key claims: the dual-use worry about automated safety research is overblown; what research sabotage would actually look like; and why "stop if we catch the AIs scheming" is not a plan you can rely on.
Companion write-ups include an actionable list of AI safety research we can do today, and a piece separating the four different meanings of "automated alignment research". He also tracks follow-up work, e.g. Jan Leike's January 2026 report "Alignment is not solved" with automated auditing scores across three Claude releases.
More from Safety
- Exabeam exec: hardest AI security problems now live outside the model — virtualsteve · 2026-09-22
- Automated reinforcement learning should scare you: from AlphaGo to math to bio labs — hattusili-the-third · 2026-09-22
- Stanford Accused of Using AI to Alter Students' Race and Gender in Ads — Polymarket · 2026-09-22
- ChatGPT reportedly refuses simple questions unless users grant email access — RexDouglass · 2026-09-22
- OpenAI calls for US leadership in setting global AI standards — Anxious-Yoghurt-9207 · 2026-09-22
- Forging 1024-bit RSA signatures in nearly SNFS time, sans factoring N — matthew_d_green · 2026-09-22