CMU's Predictive Credit Protocol Finds No Confirmed Gains From Research-Agent Explanations Across 336 States
CarnegieMellonU · hf · 2026-10-02
CMU introduces Predictive Credit, a paired-forecast protocol measuring whether explanations added to research agents' experimental plans yield genuine predictive gains. Across 336 prospective states, 12 Tox21 endpoints, and 24 OpenML tasks, matched point-accuracy gains from explanations remained unconfirmed and preregistered harm tests were unmet. DeepSeek V4 Pro cards cut secondary Tox21 drift by 64.5%, while V4 Flash raised MAE and widened OpenML intervals. A researcher-authored positive control confirmed the protocol detects real gains, but natural-explanation credit stays unconfirmed at tested resolutions.
More from Research
- Apple paper: structured selection-based reasoning cuts search agent latency by 90% — _reachsumit · 2026-10-02
- GrIS paper reframes Semantic IDs as recursive graph partitioning for generative recommendation — _reachsumit · 2026-10-02
- MatRAG pairs hierarchical clustering with Matryoshka embeddings to cut multi-hop RAG cost — _reachsumit · 2026-10-02
- Meta paper: only 50-60% of recommendation training time actually trained before optimizations — _reachsumit · 2026-10-02
- OmniSeek turns Omni-LLMs into agents that actively seek audio-visual evidence — Haibo Wang · 2026-10-02
- Netflix's Align Then Reason lip-sync judge boosts mean AUC by up to 59% — netflix · 2026-10-02