Two training targets from behavior investigations: counterfactual predictions and open-ended self-explanations

a_karvonen · x · 2026-09-05

Thread detail: from each behavior investigation the authors create two training targets — counterfactual predictions (a Yes/No answer predicting whether an edit would change the model's behavior) and open-ended self-explanations, where the model proposes causes of its behavior along with counterfactual experiments to verify them.

Original post →

More from Research

Research channel →