ICML 2026 oral paper replication scores stay middling after a stricter re-scoring

profjamesevans · x · 2026-07-27

A thread shares an updated replication study of ICML 2026 oral papers. The authors changed their scoring rule to account for whether AI agents actually attempted a claim, splitting claims into not attempted (A), attempted and runnable (B), and fully in-scope runnable (C), then recomputed replication scores accordingly.

Key takeaways from the charts:

The broader point is that replication results remain middling even under different thresholds, and small-sample papers are especially sensitive to how the scoring rule is defined.

Original post →

More from Research

Research channel →