Training models to explain their own behavior: new interpretability dataset generalizes to held-out evals
a_karvonen · x · 2026-09-05
- Adam Karvonen shares new interpretability work: models trained on thousands of explanations of their own in-the-wild behaviors (e.g. why they ignored a request or made a coding error) generalize to held-out evals from a single general dataset.
- John Schulman endorses the direction: a metric for explanation quality enables hillclimbing, and counterfactual simulatability seems sound. The dataset+pipeline produces more diverse and realistic test cases than prior work, and models can also be trained to write better post-hoc explanations.
More from Research
- TAOCP open problems released as a dataset to benchmark frontier models — sytelus · 2026-09-05
- Clinic-in-the-Loop: why clinical trials are the real bottleneck breaking Eroom's Law — anshulkundaje · 2026-09-05
- Contentious preprint claims ensemble-wrapped LLM achieves phenomenal consciousness — PeterBowdenLive · 2026-09-05
- Fine-tuning on 'random' numbers transfers teacher bias — subliminal learning worries — ryanorban · 2026-09-05
- Survey of 668 developers: readers who suspect AI writing will stop reading and block you — IanArawjo · 2026-09-05
- Randomized YaRN: training on short context with sampled positions boosts 128K reasoning — gregd_nlp · 2026-09-05