New research shows LLM leaderboards are less stable than you'd hope
beirmug · x · 2026-09-25
A paper on LLM leaderboard stability won an oral at the CTB workshop @ ICML and is now accepted to NeurIPS 2026. The authors built a unified framework for auditing how small changes to pairwise votes ripple through leaderboard rankings, finding leaderboards are far less stable than hoped. Paper and project page are public.
More from Models
- Polymarket gives SSI 53% odds of releasing its first public model by October 31 — Polymarket · 2026-09-25
- Opus 5.5 designs a buildable life-size LEGO duck: 1,113 parts, zero collisions — fofrAI · 2026-09-25
- Claude Opus 5.5 shows progress on 'sparks of AGI'-style task, higher tiers teased — iruletheworldmo · 2026-09-25
- AI pundit Tibo goes quiet after Opus 5.5 crushes rivals he said Anthropic couldn't serve — rexplosive · 2026-09-25
- User says Meta's Muse is good enough to 'retire' his self-built OpenClaw setup — analisereal · 2026-09-25
- Gated DeltaNet-2 accepted to NeurIPS 2026, claims new SOTA for hybrid linear attention — AccBalanced · 2026-09-25