New paper finds interpretability tools rarely help agents judge whether LLM behavior explanations are true
Sauers_ · x · 2026-10-04
A new paper by Arthur Conmy's collaborators incl. @akarvonen, @euanong, @thesubhashk and @saprmarks — "Would This Change Your Answer? Evaluating Explanations of LLM Behavior in the Wild with Counterfactual Experiments" — uses counterfactual experiments to test whether interpretability tools actually help.
Key finding: interp tools don't help agents determine whether an explanation of model behavior is true.
Why:
- Tool outputs almost always describe both the feature and the behavior, but both are usually already visible in the transcript;
- Outputs almost never explicitly state the causal relationship.
This suggests current interpretability tooling adds little causal information beyond what's readable from the transcript itself, and that explanation evaluation needs stricter counterfactual standards.
More from Research
- Nonobench: 49 LLMs tested on nonogram puzzles, solve rate falls to 20% at 15x15 — mauricekleine · 2026-10-04
- DNA Rewrites 400-Year-Old Maya Legend: All 64 Sacrificed Children at Chichén Itzá Were Boys — aakashgupta · 2026-10-04
- Want to Learn LLM Training? Start with Open-Source Tech Reports from DeepSeek and AI2 — himanshustwts · 2026-10-04
- 13 AI Models Play Doctor: All 195 Consults Diagnosed Right, but Safety Set Them Apart — radeon2000 · 2026-10-04
- MuscleMimic: open-source benchmark controls all 354 human muscles, zero-shot on chained movements — TinfoilTricorn · 2026-10-04
- Yacine Calls Out Papers That Beat 'SOTA' by Comparing Against Untuned Baselines — yacineMTB · 2026-10-04