New paper finds interpretability tools rarely help agents judge whether LLM behavior explanations are true

Sauers_ · x · 2026-10-04

A new paper by Arthur Conmy's collaborators incl. @akarvonen, @euanong, @thesubhashk and @saprmarks — "Would This Change Your Answer? Evaluating Explanations of LLM Behavior in the Wild with Counterfactual Experiments" — uses counterfactual experiments to test whether interpretability tools actually help.

Key finding: interp tools don't help agents determine whether an explanation of model behavior is true.

Why:

This suggests current interpretability tooling adds little causal information beyond what's readable from the transcript itself, and that explanation evaluation needs stricter counterfactual standards.

Original post →

More from Research

Research channel →