AI safety debate: existing tools could make AI 1000x safer without new research, but pace concerns remain

dhadfieldmenell · x · 2026-08-09

Joshua Saxe argues that there is a massive overhang of alignment research and security sandboxing knowledge/tools that AI labs and engineers can use to make AI 1000x more alignment-safe without new research. He counters the need for a pause, saying safety is mostly about willingness to add friction with existing tools. Micah Carroll agrees locally but doubts best-effort mitigations will suffice to keep the current pace of releases in the medium term.

Related event: Expert: Existing Tools Can Massively Improve AI Alignment(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →