CMU's Predictive Credit Protocol Finds No Confirmed Gains From Research-Agent Explanations Across 336 States

CarnegieMellonU · hf · 2026-10-02

CMU introduces Predictive Credit, a paired-forecast protocol measuring whether explanations added to research agents' experimental plans yield genuine predictive gains. Across 336 prospective states, 12 Tox21 endpoints, and 24 OpenML tasks, matched point-accuracy gains from explanations remained unconfirmed and preregistered harm tests were unmet. DeepSeek V4 Pro cards cut secondary Tox21 drift by 64.5%, while V4 Flash raised MAE and widened OpenML intervals. A researcher-authored positive control confirmed the protocol detects real gains, but natural-explanation credit stays unconfirmed at tested resolutions.

Original post →

More from Research

Research channel →