Why did the model do this? Interp tools can answer it, argues Dan Balsam

burny_tech · x · 2026-09-05

In the same thread, Dan Balsam clarifies his claim: for arbitrary questions about a model — why it did something or how to prevent it — interpretability tools combined with agents can yield meaningfully complete answers. Models are ultimately programs, and answers are knowable.

Related event: Interp Tools Plus Agents May Fully Answer Why Models Behave As They Do(2 posts)→

Original post →

More from Models

Models channel →