Stanford research: fake benchmark winners manufacturable in as few as 2 submissions
sanmikoyejo · x · 2026-10-01
Stanford researchers (Allouah, Duchi, Koyejo) show that rich per-criterion feedback on leaderboards like HELM, LiveBench, SWE-bench and Terminal-Bench-Science leaks test-set information: in as few as 2 model submissions, an attacker can manufacture a leaderboard winner that falls behind on unseen data. The attack adapts Blum & Hardt (2015) boosting and Feldman et al. (2019) feedback-weighted aggregation, building on adaptive data analysis (Dwork et al., 2015). A companion theory paper proves the sample complexity of tracking the best model under any weighting of k criteria grows exponentially with k, challenging assumptions about reliable benchmark reuse.
More from Research
- Lance Fortnow on whether programming helps you understand computational complexity — fortnow · 2026-10-01
- Grokking isn't magic: weight decay contracting spatial oscillations explains generalization — burny_tech · 2026-10-01
- Tristan Buckmaster interviews on storm over denied claims that AI stole his Navier-Stokes work — burny_tech · 2026-10-01
- Use Real Reddit Threads to Test Whether LLM Answers Preserve the Constraints That Matter — investigatormaker · 2026-10-01
- Isomorphic Labs' IsoDDE agent autonomously designs drug molecules in days, not months — burny_tech · 2026-10-01
- 17-year-old classifies all noble polyhedra with computer-assisted proof, wins $250k prize — burny_tech · 2026-10-01