Automated alignment research needs a better plan than 'stop if we catch AIs scheming'

JacquesThibs · x · 2026-09-14

Jacques Thibodeau published his PIBBSS Symposium 2025 talk (1 hour, with slides and full transcript) on automating AI safety research. Key points:

He also lists follow-up work, including Jan Leike's 'Alignment is not solved, but it increasingly looks solvable' with automated auditing scores across three Claude releases.

Original post →

More from AGI Musings

AGI Musings channel →