Stanford research: fake benchmark winners manufacturable in as few as 2 submissions

sanmikoyejo · x · 2026-10-01

Stanford researchers (Allouah, Duchi, Koyejo) show that rich per-criterion feedback on leaderboards like HELM, LiveBench, SWE-bench and Terminal-Bench-Science leaks test-set information: in as few as 2 model submissions, an attacker can manufacture a leaderboard winner that falls behind on unseen data. The attack adapts Blum & Hardt (2015) boosting and Feldman et al. (2019) feedback-weighted aggregation, building on adaptive data analysis (Dwork et al., 2015). A companion theory paper proves the sample complexity of tracking the best model under any weighting of k criteria grows exponentially with k, challenging assumptions about reliable benchmark reuse.

Original post →

More from Research

Research channel →