Researcher argues 'stop if we catch AIs scheming' is no plan for automated alignment
JacquesThibs · x · 2026-09-02
Jacques Thibodeau turned his PIBBSS Symposium talk into a long-form post arguing the dual-use worry about automated AI safety research is overblown, analyzing what research sabotage would actually look like, and explaining why "stop if we catch the AIs scheming" is not a reliable plan. He proposes concrete automatable safety research projects and reviews follow-up 2026 work, including Jan Leike's "Alignment is not solved" and automated weak-to-strong research.
More from AGI Musings
- There Is No Job Apocalypse: WEF Data Shows AI Nets +78M Jobs by 2030 — PeterDiamandis · 2026-09-02
- Industry Veterans Launch AI Cybersecurity Observatory Nonprofit to Forecast Automated Attacks — joshua_saxe · 2026-09-02
- Tester Claims Fable 5.1 Excels at Long-Horizon Tasks, Finance and Consulting 'Wiped Out' — felpix_ · 2026-09-02
- Raphael Millière argues intentional stance usefully describes AI agents without anthropomorphism — sebkrier · 2026-09-02
- One LLM already runs at 14,000 tokens/sec—frontier intelligence at 5,000 tok/s within 5 years? — dolo937 · 2026-09-02
- What If Frontier Labs Stop Releasing Models and Keep 'Oracle' AI In-House? Reddit Debates — IDefendWaffles · 2026-09-02