CHIVE: Evaluating LLM explanations via counterfactual experiments — common interpretability techniques show no uplift
a_karvonen · x · 2026-08-22
A new paper from the Anthropic Fellows Program, "Would this change your answer?", proposes evaluating explanations via counterfactual simulatability: whether an explanation helps predict model behavior on related counterfactual inputs.
The authors introduce CHIVE (Counterfactual Hypothesis Investigation Via Edits), an agentic pipeline that discovers unexpected model behaviors in the wild and investigates them with counterfactual prompt edits, yielding thousands of explanations with supporting counterfactual evidence. Two key findings:
- When evaluating common LLM interpretability techniques, they surprisingly find no uplift in an agent's ability to predict counterfactual model behavior from any technique studied.
- Using CHIVE investigations as training data — teaching models to predict outcomes of counterfactual experiments — generalizes to held-out OOD settings (e.g., detecting when a hint changed an answer), though results vary with training format.
More from Safety
- Expert Advocates Human-in-the-Loop for Agentic Security in Critical Software — thedealdirector · 2026-08-22
- AI in HR is a legal and ethical minefield: Meta lawsuit as a warning — DavidLinthicum · 2026-08-22
- Security Experts Warn AI Agents Can Mount Sophisticated Human Deception — pstAsiatech · 2026-08-22
- Reddit Cuts Hate Content Enforcement Time to Under 5 Seconds Using AI — lilyraynyc · 2026-08-22
- Anthropic deploys Claude Mythos 5 for cyber defense via security scanner — The Decoder · 2026-08-22
- AI Text Watermarking Is Free And Good, Explained by Aaronson — TheZvi · 2026-08-22