RECAP lets probes verify activation explanations that reconstruction scores can fake

Hiskias Dingeto · hf · 2026-07-23

RECAP makes activation explanations independently checkable instead of just reconstructable

This paper argues that high reconstruction scores for natural-language explanations of hidden activations do not guarantee truthful claims. The standard test can be satisfied by gist matching or by private co-adapted codes that game reconstruction without being faithful.

The authors demonstrate the problem in two settings:

To address this, they introduce:

Reported results include:

The takeaway: reconstruction alone is not enough to certify explanation faithfulness; separate probes can make internal content independently checkable.

Original post →

More from Research

Research channel →