Anthropic Fellows train models to explain their own wild behaviors with generalization to held-out evals

a_karvonen · x · 2026-09-05

Adam Karvonen and collaborators (Anthropic Fellows Program) asked whether models can learn to explain their own behaviors, such as ignoring a user request or making a coding error. Using a CHIVE pipeline that finds unexpected in-the-wild behaviors and probes them with counterfactual prompt edits, they produced thousands of high-quality explanations and trained models on them. Training on this single general dataset generalizes to held-out evaluations.

Related event: Anthropic Fellows Release CHIVE for Model Self-Explanation(2 posts)→

Original post →

More from Research

Research channel →