Two training targets from behavior investigations: counterfactual predictions and open-ended self-explanations
a_karvonen · x · 2026-09-05
Thread detail: from each behavior investigation the authors create two training targets — counterfactual predictions (a Yes/No answer predicting whether an edit would change the model's behavior) and open-ended self-explanations, where the model proposes causes of its behavior along with counterfactual experiments to verify them.
More from Research
- Delaunay Canopy (ECCV 2026): SOTA building wireframe reconstruction from sparse LiDAR — RexDouglass · 2026-09-05
- Researchers open Postdoc/PhD role on privacy and contextual integrity in AI agents — niloofar_mire · 2026-09-05
- One bit is enough: hidden data encoded via AC polarity flips across systems — MoonL88537 · 2026-09-05
- RNASSTR: new Rfam-based RNA secondary structure dataset with structure-aware train/test splits — chaitjo · 2026-09-05
- Self-explanation training generalizes beyond narrow hint formats to held-out evals — a_karvonen · 2026-09-05
- Anthropic Fellows train models to explain their own wild behaviors with generalization to held-out evals — a_karvonen · 2026-09-05