New research shows LLM leaderboards are less stable than you'd hope

beirmug · x · 2026-09-25

A paper on LLM leaderboard stability won an oral at the CTB workshop @ ICML and is now accepted to NeurIPS 2026. The authors built a unified framework for auditing how small changes to pairwise votes ripple through leaderboard rankings, finding leaderboards are far less stable than hoped. Paper and project page are public.

Original post →

More from Models

Models channel →