Training models to explain their own behavior: new interpretability dataset generalizes to held-out evals

a_karvonen · x · 2026-09-05

Related event: Anthropic Fellows Train LLMs to Explain Their Own Behavior, Generalizing to Unseen Evals(5 posts)→

Original post →

More from Research

Research channel →