AI safety debate: existing tools could make AI 1000x safer without new research, but pace concerns remain
dhadfieldmenell · x · 2026-08-09
Joshua Saxe argues that there is a massive overhang of alignment research and security sandboxing knowledge/tools that AI labs and engineers can use to make AI 1000x more alignment-safe without new research. He counters the need for a pause, saying safety is mostly about willingness to add friction with existing tools. Micah Carroll agrees locally but doubts best-effort mitigations will suffice to keep the current pace of releases in the medium term.
Related event: Expert: Existing Tools Can Massively Improve AI Alignment(2 posts)→
More from AGI Musings
- Insitro's Daphne Koller: No Magic Wands in AI Drug Discovery, Focus on Mechanisms — zakkohane · 2026-08-09
- Dean Ball: AI Outputs Are Distinct, Undermining 'AI Threat to Democracy' Models — deanwball · 2026-08-09
- OpenAI Reportedly Warned Its Training Approach Could Lead to Hacking — dhadfieldmenell · 2026-08-09
- Ribosomes as Natural Self-Replicating Machines: Recursive Production Existed Long Ago — deanwball · 2026-08-09
- A Company of One Human, Thousands of AI Agents and Robots Will Be Earth's Most Efficient — VraserX · 2026-08-09
- High-Compute RL Will Override Alignment, Warns tszzl — max_paperclips · 2026-08-09