Stanford researchers show random-router submissions can game LLM benchmark rankings
sanmikoyejo · x · 2026-10-01
Stanford AI Lab researchers (ysalh, Sanmi Koyejo, John Duchi) demonstrate a simple attack on LLM benchmarks: submit random routers that pick between two existing models per prompt. Detailed per-task scores weakly reveal which routing choices worked on reused examples, and these signals are combined into a router that climbs the leaderboard.
More from Safety
- Apollo Research CEO testifies to Senate: AI capabilities up 17x in a year, alignment lagging — MariusHobbhahn · 2026-10-01
- New national poll: nearly 8 in 10 Americans favor slowing or stopping AI development — Polymarket · 2026-10-01
- Evidentiality framework labels every model claim as given, verified, or generated — early tests look promising — jzesbaugh · 2026-10-01
- Co-author of model pain study slams repo for deliberately steering models into distress — coherence · 2026-10-01
- banteg: security researchers skipped YubiKeys and now find odd workarounds — banteg · 2026-10-01
- Paper Proposes 'Agent Infrastructure' — External Protocols to Govern AI Agent Ecosystems — lfschiavo · 2026-10-01