NeurIPS paper: sub-1% targeted perturbations can flip LMArena's top-ranked model
lintool · x · 2026-09-26
- New NeurIPS-accepted paper introduces a unified perturbation framework for Bradley-Terry leaderboards like LMArena, using influence-based approximations
- Studies three match-level perturbations (Drop, Add, Flip) plus player removal, measuring effects on top-k membership, Kendall's tau ranking consistency, and confidence intervals
- Across Chatbot Arena and six other pairwise datasets, leaderboards are non-robust on all three objectives: sub-1% targeted perturbations can change the top-ranked model
- The same influence scores enable efficient targeted manipulation, promoting/demoting specific models with fewer actions than prior manipulation and active-sampling baselines
- Provides normalized dataset-level robustness scores as a practical auditing tool and motivation for sturdier evaluation protocols
More from Research
- Anthropic says Claude can compute notoriously hard Nine Loops physics amplitudes — daniel_mac8 · 2026-09-26
- Researchers teach LLMs to find interesting theorems, boosting interestingness 4.3x — CatAstro_Piyush · 2026-09-26
- Anthropic: Claude solves nine-loop scattering amplitudes, breaking the eight-loop record — AnthropicAI · 2026-09-26
- Both "grep is all you need" and "BM25 is all you need" Papers Just Got Accepted — lintool · 2026-09-26
- Reasonable Team Publishes TLA+ Tutorial: Not a Silver Bullet, AI Agents Could Change That — fhuszar · 2026-09-26
- Bilevel optimization densifies scarce labels to fix OOD molecular property prediction (ICML WS 2025) — CatAstro_Piyush · 2026-09-26