Security Expert: Existing Tools Can Make AI 1000x More Aligned, No Research Pause Needed

joshua_saxe · x · 2026-08-09

Security expert Joshua Saxe argues that there is a massive overhang in alignment research and security sandboxing knowledge/tools just waiting to be adopted by AI labs and application engineers.

He contends that utilizing these existing tools would make AI 1000x more alignment-safe without requiring any new research. Pushing back against narratives calling for a Los Alamos-style pause to solve alignment, Saxe emphasizes that safety isn't primarily a question of new knowledge right now; rather, it's a matter of willingness to dial in more friction using the tools we already have.

Original post →

More from AGI Musings

AGI Musings channel →