AI safety researcher: interpretability and alignment methods work but not reliably, and we don't know how to fix that
DavidSKrueger · x · 2026-09-22
David Krueger, AI safety researcher, responded to skeptics by noting he has published work in all four relevant areas — interpretability, testing, alignment, and control — at top AI venues. His point: the methods exist and are not useless, but they don't work reliably, and the field doesn't know how to fix that. We can't understand how AI systems work, predict their behavior, prevent misbehavior, or stay in control when it happens — foundational open problems despite years of effort.
Related event: Cambridge Researcher: Four Unresolved Problems Make AI Safety No Quick Fix(2 posts)→
More from AGI Musings
- AI's constant caveats and follow-up questions are making researchers think harder — geoffwolfe · 2026-09-22
- Researcher Suspects AI Desk-Rejection Tools Create More Work Than They Save — IanArawjo · 2026-09-22
- Vinod Khosla: Personal AI moats are data trust and getting tasks done — brucemacv · 2026-09-22
- Halvar Flake: CS Is Turning From an Exact Field Into a Probabilistic One — soumitrashukla9 · 2026-09-22
- Epoch AI researchers: 2026 is the first year AI capabilities meaningfully hit the world — Jsevillamol · 2026-09-22
- Slop papers leave a public trail: reviewers remember, and it can sink your research career — cocoweixu · 2026-09-22