Anthropic Fellows train models to explain their own wild behaviors with generalization to held-out evals
a_karvonen · x · 2026-09-05
Adam Karvonen and collaborators (Anthropic Fellows Program) asked whether models can learn to explain their own behaviors, such as ignoring a user request or making a coding error. Using a CHIVE pipeline that finds unexpected in-the-wild behaviors and probes them with counterfactual prompt edits, they produced thousands of high-quality explanations and trained models on them. Training on this single general dataset generalizes to held-out evaluations.
Related event: Anthropic Fellows Release CHIVE for Model Self-Explanation(2 posts)→
More from Research
- Delaunay Canopy (ECCV 2026): SOTA building wireframe reconstruction from sparse LiDAR — RexDouglass · 2026-09-05
- Researchers open Postdoc/PhD role on privacy and contextual integrity in AI agents — niloofar_mire · 2026-09-05
- One bit is enough: hidden data encoded via AC polarity flips across systems — MoonL88537 · 2026-09-05
- RNASSTR: new Rfam-based RNA secondary structure dataset with structure-aware train/test splits — chaitjo · 2026-09-05
- Self-explanation training generalizes beyond narrow hint formats to held-out evals — a_karvonen · 2026-09-05
- Two training targets from behavior investigations: counterfactual predictions and open-ended self-explanations — a_karvonen · 2026-09-05