Self-explanation training generalizes beyond narrow hint formats to held-out evals

a_karvonen · x · 2026-09-05

Prior self-explanation training generalized narrowly (e.g., one hint format to another) due to narrow datasets. The authors' training never targeted hints, yet both models improve on the held-out hint eval ("would removing the hint change your answer?") and on other evals.

Related event: Self-Explanation Training Generalizes Across Models and Tasks(2 posts)→

Original post →

More from Research

Research channel →