Elo Failed for Model Ranking, So I Used Migration-Frequency Normalization
cloneofsimo · x · 2026-10-06
The author shares lessons from building a model leaderboard scoring system: Elo didn't work, and neither did other normalized ranking models.
The final approach is based on user migration frequency:
- Let Fij be the number of people moving from model i to j
- Tij = Fij / sumj Fij (i-to-j normalized per i)
- Sj = sumj Tij, i.e. when people leave a model, how often they choose j
So ranking is driven by where users go when they switch, not head-to-head wins.
More from Models
- SelfBench turns real GitHub PRs into evals: open-weight models cost more and do worse — ycombinator · 2026-10-06
- Embedded Grok on X reportedly lacks per-user context isolation, called out as a major flaw — altryne · 2026-10-06
- Liquid AI's d1 decision model adds vision, beats GPT-6.1 on 4 of 6 tasks at up to 200x lower cost — JosephJacks_ · 2026-10-06
- Anthropic reviewers alerted police to a Claude chat threatening a sheriff's office, leading to an arrest — rohanpaul_ai · 2026-10-06
- flow-1: RL-trained model matches GPT-6-sol at trace debugging while 23x cheaper — kalyan_kpl · 2026-10-06
- First large-scale 3B/8B continuous diffusion LMs match pass@1 and beat pass@k vs masked dLMs — ArashVahdat · 2026-10-06