AI safety researcher argues against "solving alignment" as the field's core frame
joshua_saxe · x · 2026-09-28
AI researcher Joshua Saxe published a long-form argument against framing AI safety around "solving alignment" and "solving interpretability":
- Civilization already runs on approximate control: modern society rests on imperfectly forecasting and steering chaotic natural and social systems — climate, biology, global supply chains no one fully understands. AI systems are becoming similarly opaque "God-like" systems we will never fully comprehend, and will understand less as capabilities advance.
- The interpretability vision lacks evidence: the prevailing narrative in policy circles, labs, and media holds that we'll eventually understand model internals with many nines of precision and read out guarantees models won't sabotage critical systems. But after 14+ years, interpretability has produced nothing like the reliable, scalable understanding that vision requires — and the gap is widening. He contrasts carefully-tuned t-SNE plots in safety papers with capability news of models solving verifiable Millennium Prize problems, noting the epistemic hollowness of the former.
- Not anti-interpretability: it's valuable as data exploration and for building wrong-but-useful theories, much like crude economic models give policy discussions a vocabulary — but there's no evidence it will yield formal proofs or fine-grained control.
- The real goal: building a civilization that can safely live with layered multi-agent systems we can never make provable claims about. AI safety should borrow from industrial safety practice and disaster planning for hurricanes, nuclear safety, pandemics, and financial crises — assuming incomplete world models, forecasting imperfectly, and monitoring what actually happens.
More from AGI Musings
- Indie dev on AI: your game gets cloned within days and slop buries originals — ring_hyacinth · 2026-09-28
- AI Personal Assistants May Quietly Replace Much of Normal Human Interaction — GarrisonLovely · 2026-09-28
- Paras Chopra: AI's Pull Comes From Looming Human Economic Irrelevance — paraschopra · 2026-09-28
- Jan Kulveit: Stop Overcorrecting Toward 'Power-Seeker' Readings of Frontier Models — jankulveit · 2026-09-28
- Silicon Valley Parenting: Kids' Best Years Consumed by Sports and Accelerated Classes — vaibhavbetter · 2026-09-28
- Hilbert Spaess Argues a Global AI Slowdown Serves Everyone's Interest — hilbertspaess · 2026-09-28