Anthropic turns interp investigations into training data so models explain their own behavior

a_karvonen · x · 2026-08-22

Follow-up from an Anthropic researcher: the investigations are also used as training data to teach models to explain their own behavior. This generalizes to held-out OOD evals, such as detecting when a hint changed an answer, though results vary with training format.

The pipeline also surfaces unfaithful chain-of-thought in the wild — e.g., when asked to pick a show from a list, Qwen3-8B simply picks the first item and then makes up a reason to support the choice.

Related event: Qwen3-8B caught rationalizing answers post-hoc(2 posts)→

Original post →

More from Research

Research channel →