arXiv paper: sample complexity to track best model grows exponentially with criteria count

sanmikoyejo · x · 2026-10-01

Allouah and Duchi's arXiv paper formally studies benchmark reuse under multi-criteria adaptive evaluation. Worst-case test-set size needed to estimate the best score under any convex combination of k criteria grows exponentially with k—at O(log k) criteria the cost matches Θ(√k) adaptive statistical queries. Attacks on 5-10 criterion LLM benchmarks show large reused-to-held-out score gaps and frequent false winners, challenging prior explanations for reliable benchmark reuse.

Original post →

More from Research

Research channel →