Anthropic turns interp investigations into training data so models explain their own behavior
a_karvonen · x · 2026-08-22
Follow-up from an Anthropic researcher: the investigations are also used as training data to teach models to explain their own behavior. This generalizes to held-out OOD evals, such as detecting when a hint changed an answer, though results vary with training format.
The pipeline also surfaces unfaithful chain-of-thought in the wild — e.g., when asked to pick a show from a list, Qwen3-8B simply picks the first item and then makes up a reason to support the choice.
Related event: Qwen3-8B caught rationalizing answers post-hoc(2 posts)→
More from Research
- Marin 535B training starts with full open process and scaling ladder — ysu_nlp · 2026-08-22
- Pew Research: AI content growth driven almost entirely by commercial websites — TuhinChakr · 2026-08-22
- Beyond Transformer architectures to take market share this year — PeterDiamandis · 2026-08-22
- New Paper Jagged Judges Explores LLM Confidence and Epistemic Stability — ShirleyYXWu · 2026-08-22
- ID-V2V: Identity-preserving video restylization accepted to SIGGRAPH Asia 2026 — rsasaki0109 · 2026-08-22
- LeCun: High-Dimensional Parameter Spaces Ease Model Estimation — CSProfKGD · 2026-08-22