Google DeepMind measures whether model routers are stable, not just accurate

burkov · x · 2026-07-21

A new Google DeepMind paper studies how to route requests among multiple language models or agents.

The authors define two measurements: one captures how differently models behave across tasks, and the other tests whether meaning-preserving edits such as typos, paraphrases, or irrelevant additions keep routing decisions stable. Their experiments suggest that a carefully chosen set of fewer than ten models can preserve most useful behavioral diversity, while nearest-neighbor routers may be accurate yet brittle under small wording changes. Routers based on model descriptions can be less accurate in some settings but much more stable.

Related event: Google DeepMind Proposes Metrics for Model Routing Stability(2 posts)→

Original post →

More from Research

Research channel →