CHIVE: Evaluating LLM explanations via counterfactual experiments — common interpretability techniques show no uplift

a_karvonen · x · 2026-08-22

A new paper from the Anthropic Fellows Program, "Would this change your answer?", proposes evaluating explanations via counterfactual simulatability: whether an explanation helps predict model behavior on related counterfactual inputs.

The authors introduce CHIVE (Counterfactual Hypothesis Investigation Via Edits), an agentic pipeline that discovers unexpected model behaviors in the wild and investigates them with counterfactual prompt edits, yielding thousands of explanations with supporting counterfactual evidence. Two key findings:

Related event: Study Finds Common Interpretability Tools Offer No Gain in LLM Behavior Evaluation(2 posts)→

Original post →

More from Safety

Safety channel →