Self-explanation training generalizes beyond narrow hint formats to held-out evals
a_karvonen · x · 2026-09-05
Prior self-explanation training generalized narrowly (e.g., one hint format to another) due to narrow datasets. The authors' training never targeted hints, yet both models improve on the held-out hint eval ("would removing the hint change your answer?") and on other evals.
Related event: Self-Explanation Training Generalizes Across Models and Tasks(2 posts)→
More from Research
- Delaunay Canopy (ECCV 2026): SOTA building wireframe reconstruction from sparse LiDAR — RexDouglass · 2026-09-05
- Researchers open Postdoc/PhD role on privacy and contextual integrity in AI agents — niloofar_mire · 2026-09-05
- One bit is enough: hidden data encoded via AC polarity flips across systems — MoonL88537 · 2026-09-05
- RNASSTR: new Rfam-based RNA secondary structure dataset with structure-aware train/test splits — chaitjo · 2026-09-05
- Two training targets from behavior investigations: counterfactual predictions and open-ended self-explanations — a_karvonen · 2026-09-05
- Anthropic Fellows train models to explain their own wild behaviors with generalization to held-out evals — a_karvonen · 2026-09-05