PIBBSS talk: "Stop if we catch AIs scheming" is no plan for automated alignment

JacquesThibs · x · 2026-09-22

Jacques Thibodeau published the full written and recorded version (63 min) of his September 2025 PIBBSS Symposium talk on automating AI safety research. Key claims: the dual-use worry about automated safety research is overblown; what research sabotage would actually look like; and why "stop if we catch the AIs scheming" is not a plan you can rely on.

Companion write-ups include an actionable list of AI safety research we can do today, and a piece separating the four different meanings of "automated alignment research". He also tracks follow-up work, e.g. Jan Leike's January 2026 report "Alignment is not solved" with automated auditing scores across three Claude releases.

Original post →

More from Safety

Safety channel →