Interp tools plus agents can give meaningfully complete answers to why models do what they do
burny_tech · x · 2026-09-05
Dan Balsam argues that for arbitrary questions like "why did the model do this?" or "how can I stop it?", increasingly capable interpretability tools — used by increasingly capable agents — can produce meaningfully complete answers. Models are programs, he notes; not holding every line in your head doesn't make the answers unknowable.
Related event: Interp Tools Plus Agents May Fully Answer Why Models Behave As They Do(2 posts)→
More from Models
- GPT progression: GPT-3 text, GPT-4 images, GPT-5 video, GPT-6 3D? — PaulYacoubian · 2026-09-05
- Open models should chase depth, not Claude-style coding, argues researcher — teortaxesTex · 2026-09-05
- Reddit user claims Astra's capability jump exceeds GPT-3.5→4, yet the internet is silent — Glittering-Neck-2505 · 2026-09-05
- GPT-6 Astra claimed as new SOTA in agentic CAD — mckbrando · 2026-09-05
- GPT-6 hyped as AGI can't even render a snake with its head attached — SnooCheesecakes1893 · 2026-09-05
- GPT-6 Astra usage limits hit every 5 minutes, user can't even reset anymore — memeka · 2026-09-05