Automated alignment research needs a better plan than 'stop if we catch them scheming'

JacquesThibs · x · 2026-09-11

Jacques Thibodeau published his PIBBSS Symposium talk (Sept 2025, 63 min) with slides and transcript. Key claims: the dual-use worry about automated safety research is overblown; he analyzes what research sabotage would actually look like; and 'stop if we catch the AIs scheming' is not a plan you can rely on. The post links his earlier essays on automating AI safety today and on the four meanings of 'automated alignment research', plus follow-up work like Jan Leike's 'Alignment is not solved' (Jan 2026).

Original post →

More from AGI Musings

AGI Musings channel →