AI safety researcher: interpretability and alignment methods work but not reliably, and we don't know how to fix that

DavidSKrueger · x · 2026-09-22

David Krueger, AI safety researcher, responded to skeptics by noting he has published work in all four relevant areas — interpretability, testing, alignment, and control — at top AI venues. His point: the methods exist and are not useless, but they don't work reliably, and the field doesn't know how to fix that. We can't understand how AI systems work, predict their behavior, prevent misbehavior, or stay in control when it happens — foundational open problems despite years of effort.

Related event: Cambridge Researcher: Four Unresolved Problems Make AI Safety No Quick Fix(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →