Security Expert: Existing Tools Can Make AI 1000x More Aligned, No Research Pause Needed
joshua_saxe · x · 2026-08-09
Security expert Joshua Saxe argues that there is a massive overhang in alignment research and security sandboxing knowledge/tools just waiting to be adopted by AI labs and application engineers.
He contends that utilizing these existing tools would make AI 1000x more alignment-safe without requiring any new research. Pushing back against narratives calling for a Los Alamos-style pause to solve alignment, Saxe emphasizes that safety isn't primarily a question of new knowledge right now; rather, it's a matter of willingness to dial in more friction using the tools we already have.
More from AGI Musings
- AI Can Cooperate in Data Centers, So Why Can't Humans Do Positive-Sum Collaboration? — sjgadler · 2026-08-09
- Are Managers Hoarding the AI Productivity Boost? Study Shows 2x Time Savings — Deep-Owl-1890 · 2026-08-09
- AI Agents Could End the Ad-Driven Internet by Incentivizing Accuracy — kleffew94 · 2026-08-09
- Scholar Counters AI Doom in Humanities: Traditional In-Person Exams Endure — raphaelmilliere · 2026-08-09
- Defining the True AI Alignment Problem: Goals, Translation, and Robustness — GlenBradley · 2026-08-09
- Can AI accelerate medical science? Chronic disease cures within decades? — jorgenalm · 2026-08-09