Benchmark granularity magnifies ranking attack: 18 task scores quadruple selection loss
sanmikoyejo · x · 2026-10-01
Score granularity drives the benchmark attack: on LiveBench, returning 18 per-task scores instead of one aggregate score almost quadruples the mean model-selection loss after just 8 router submissions.
More from Research
- Lance Fortnow on whether programming helps you understand computational complexity — fortnow · 2026-10-01
- Grokking isn't magic: weight decay contracting spatial oscillations explains generalization — burny_tech · 2026-10-01
- Tristan Buckmaster interviews on storm over denied claims that AI stole his Navier-Stokes work — burny_tech · 2026-10-01
- Use Real Reddit Threads to Test Whether LLM Answers Preserve the Constraints That Matter — investigatormaker · 2026-10-01
- Isomorphic Labs' IsoDDE agent autonomously designs drug molecules in days, not months — burny_tech · 2026-10-01
- 17-year-old classifies all noble polyhedra with computer-assisted proof, wins $250k prize — burny_tech · 2026-10-01