Fitting a performance cone with bootstrapped models to test if AI leaderboard rankings actually hold
PTenigma · x · 2026-09-21
- PTenigma proposes a statistical fix for unreliable model ranking challenges: train models on bootstrapped subsets of the training data, then fit a "performance cone" to detect whether rankings actually cross over as compute scales.
- Key points:
- Use Ville's law of confidence sequences to determine the fewest bootstrap draws needed while capping the risk of missing a crossover.
- For fully open-source models, paired bootstrapped datasets enable paired tests, which are typically more sensitive.
- He argues most current challenges skip this and may produce rankings that can't be relied upon.
More from Research
- GEPA Prompt Optimization Lifts Jev's F1 From 69.1% to 79.7% on Medical Literature Task — matei_zaharia · 2026-09-21
- Philosophy Needs to Become Robust RL Objectives, Not Thought Experiments — willcb · 2026-09-21
- The Illustrated Recurrence: From Amari-Hopfield Nets to GPT-6 Astra — gklambauer · 2026-09-21
- Anthropic builds Bay Area wet lab where Claude will direct robots to speed drug research — FinanceYF5 · 2026-09-21
- DraftTrace assesses student learning from process, not just product — keviv9 · 2026-09-21
- Matthew Berman asks: is this simulated fruit fly real? — Matthew Berman · 2026-09-21