Automated alignment research needs a better plan than 'stop if we catch them scheming'
JacquesThibs · x · 2026-09-11
Jacques Thibodeau published his PIBBSS Symposium talk (Sept 2025, 63 min) with slides and transcript. Key claims: the dual-use worry about automated safety research is overblown; he analyzes what research sabotage would actually look like; and 'stop if we catch the AIs scheming' is not a plan you can rely on. The post links his earlier essays on automating AI safety today and on the four meanings of 'automated alignment research', plus follow-up work like Jan Leike's 'Alignment is not solved' (Jan 2026).
More from AGI Musings
- DSPy creator endorses "bitter free lunch": methods are compute-dominated, problem specification is what matters — lateinteraction · 2026-09-11
- DSPy creator endorses 'The Bitterest Lesson' sequel: specify problems, not methods — lateinteraction · 2026-09-11
- Cambridge AI safety researcher David Krueger warns 'AI could kill us all' — DavidSKrueger · 2026-09-11
- Cambridge's David Krueger launches movement on existential AI risk, opens sign-ups — DavidSKrueger · 2026-09-11
- Recursive Self-Improvement Deemed More Plausible — and Scarier — Than AI Consciousness — AndyMasley · 2026-09-11
- How much GDP would you spend on a machine that only cures diseases? — adityaag · 2026-09-11