Researcher argues 'stop if we catch AIs scheming' is no plan for automated alignment

JacquesThibs · x · 2026-09-02

Jacques Thibodeau turned his PIBBSS Symposium talk into a long-form post arguing the dual-use worry about automated AI safety research is overblown, analyzing what research sabotage would actually look like, and explaining why "stop if we catch the AIs scheming" is not a reliable plan. He proposes concrete automatable safety research projects and reviews follow-up 2026 work, including Jan Leike's "Alignment is not solved" and automated weak-to-strong research.

Original post →

More from AGI Musings

AGI Musings channel →