Anthropic Fellows Train LLMs to Explain Their Own Behavior, Generalizing to Unseen Evals

Adam Karvonen's team (work done during the Anthropic Fellows Program, with co-authors Euan Ong, Subhash Kantamneni, and Samuel Marks) released the CHIVE paper and blog post (arXiv:2608.16747), introducing a new method for having models explain their own behavior and demonstrating that trained self-explanation ability generalizes to evals it was never trained on. John Schulman reshared it approvingly.

Confirmed

Why it matters

2026-09-05 ~ 2026-09-05 · 5 related posts

Primary sources