Stanford researchers show random-router submissions can game LLM benchmark rankings

sanmikoyejo · x · 2026-10-01

Stanford AI Lab researchers (ysalh, Sanmi Koyejo, John Duchi) demonstrate a simple attack on LLM benchmarks: submit random routers that pick between two existing models per prompt. Detailed per-task scores weakly reveal which routing choices worked on reused examples, and these signals are combined into a router that climbs the leaderboard.

Original post →

More from Safety

Safety channel →