arXiv paper: sample complexity to track best model grows exponentially with criteria count
sanmikoyejo · x · 2026-10-01
Allouah and Duchi's arXiv paper formally studies benchmark reuse under multi-criteria adaptive evaluation. Worst-case test-set size needed to estimate the best score under any convex combination of k criteria grows exponentially with k—at O(log k) criteria the cost matches Θ(√k) adaptive statistical queries. Attacks on 5-10 criterion LLM benchmarks show large reused-to-held-out score gaps and frequent false winners, challenging prior explanations for reliable benchmark reuse.
More from Research
- Lance Fortnow on whether programming helps you understand computational complexity — fortnow · 2026-10-01
- Grokking isn't magic: weight decay contracting spatial oscillations explains generalization — burny_tech · 2026-10-01
- Tristan Buckmaster interviews on storm over denied claims that AI stole his Navier-Stokes work — burny_tech · 2026-10-01
- Use Real Reddit Threads to Test Whether LLM Answers Preserve the Constraints That Matter — investigatormaker · 2026-10-01
- Isomorphic Labs' IsoDDE agent autonomously designs drug molecules in days, not months — burny_tech · 2026-10-01
- 17-year-old classifies all noble polyhedra with computer-assisted proof, wins $250k prize — burny_tech · 2026-10-01