Automated alignment research needs a better plan than 'stop if we catch AIs scheming'
JacquesThibs · x · 2026-09-14
Jacques Thibodeau published his PIBBSS Symposium 2025 talk (1 hour, with slides and full transcript) on automating AI safety research. Key points:
- The dual-use worry about automated safety research is overblown
- He separates four distinct things people mean by 'automated alignment research'
- 'Stop if we catch the AIs scheming' is not a reliable plan; he argues for something like a 'Responsible Automation Policy'
He also lists follow-up work, including Jan Leike's 'Alignment is not solved, but it increasingly looks solvable' with automated auditing scores across three Claude releases.
More from AGI Musings
- Terence Tao: LLM math is simple undergrad stuff — the real mystery is why they work — rohanpaul_ai · 2026-09-14
- Anil Seth: 'metaphors doubling back' capture AI's consciousness confusion — anilkseth · 2026-09-14
- OpenAI's NS math result reportedly took $15M of compute; efficiency questioned — rbhar90 · 2026-09-14
- "Pure resentment": viral take says AI gives high-IQ folks a taste of average life — zetalyrae · 2026-09-14
- New paper: cutting junior hiring in the AI era risks 'lost cohorts' and talent cycles — mattbeane · 2026-09-14
- Yoav Goldberg translates AI industry speak: 'third-party evaluators' are spies, 'pacing' means stop spending — yoavgo · 2026-09-14