CHIVE paper: evaluating LLM behavior explanations in the wild with counterfactual experiments
a_karvonen · x · 2026-09-05
Paper (arXiv:2608.16747) and blog post released by Adam Karvonen, Euan Ong, Subhash Kantamneni, and Samuel Marks, done as part of the Anthropic Fellows Program.
- Problem: interpretability and CoT-faithfulness research needs a way to judge what makes a good explanation of model behavior.
- Method: CHIVE (Counterfactual Hypothesis Investigation Via Edits), an agentic pipeline that finds unexpected in-the-wild model behaviors and investigates them via counterfactual prompt edits, yielding thousands of high-quality explanations with supporting evidence.
- Finding 1: judged by counterfactual simulatability, none of the common interpretability techniques studied improved an agent's ability to predict counterfactual model behaviors.
- Finding 2: training on CHIVE-generated counterfactual data generalizes to various out-of-distribution settings.
Related event: Anthropic Fellows Release CHIVE for Model Self-Explanation(2 posts)→
More from Research
- Delaunay Canopy (ECCV 2026): SOTA building wireframe reconstruction from sparse LiDAR — RexDouglass · 2026-09-05
- Researchers open Postdoc/PhD role on privacy and contextual integrity in AI agents — niloofar_mire · 2026-09-05
- One bit is enough: hidden data encoded via AC polarity flips across systems — MoonL88537 · 2026-09-05
- RNASSTR: new Rfam-based RNA secondary structure dataset with structure-aware train/test splits — chaitjo · 2026-09-05
- Self-explanation training generalizes beyond narrow hint formats to held-out evals — a_karvonen · 2026-09-05
- Two training targets from behavior investigations: counterfactual predictions and open-ended self-explanations — a_karvonen · 2026-09-05