ICML 2026 oral paper replication scores stay middling after a stricter re-scoring
profjamesevans · x · 2026-07-27
A thread shares an updated replication study of ICML 2026 oral papers. The authors changed their scoring rule to account for whether AI agents actually attempted a claim, splitting claims into not attempted (A), attempted and runnable (B), and fully in-scope runnable (C), then recomputed replication scores accordingly.
Key takeaways from the charts:
- 168 oral papers were accepted; 104 shipped runnable code.
- 105 were fully replicated by the researchers.
- Of those, 92 had at least five verifiable claims and were scored in the main analysis.
- Among those 92, only 34 reproduced over 40% of claims, and 8 reproduced over 80%.
- When the authors switched from the looser “verifiable claims” score to the stricter “fully in scope” score, the median replication score dropped from about 50% to 42% once papers had at least three judgeable claims.
The broader point is that replication results remain middling even under different thresholds, and small-sample papers are especially sensitive to how the scoring rule is defined.
More from Research
- NUS builds a soft force sensor that drives actuators without electronics or power — CurieuxExplorer · 2026-07-27
- Chelsea Finn says robot RL is bottlenecked by physical rollout cost, not algorithms — ycombinator · 2026-07-27
- Long-running agents will need immutable event logs, this thread argues — sebpaquet · 2026-07-27
- Seed IQ navigates Doom II, prompting questions about benchmarks beyond ARC-AGI — Fit_Transition8824 · 2026-07-27
- Agentic Data Science in Practice: Agents Write Code but Answer Wrong Questions — hugobowne · 2026-07-27
- A concise canon of foundational papers in ML, systems, NLP, speech, and audio — deliprao · 2026-07-27