Interp tools plus agents can give meaningfully complete answers to why models do what they do

burny_tech · x · 2026-09-05

Dan Balsam argues that for arbitrary questions like "why did the model do this?" or "how can I stop it?", increasingly capable interpretability tools — used by increasingly capable agents — can produce meaningfully complete answers. Models are programs, he notes; not holding every line in your head doesn't make the answers unknowable.

Related event: Interp Tools Plus Agents May Fully Answer Why Models Behave As They Do(2 posts)→

Original post →

More from Models

Models channel →