Google DeepMind measures whether model routers are stable, not just accurate
burkov · x · 2026-07-21
A new Google DeepMind paper studies how to route requests among multiple language models or agents.
The authors define two measurements: one captures how differently models behave across tasks, and the other tests whether meaning-preserving edits such as typos, paraphrases, or irrelevant additions keep routing decisions stable. Their experiments suggest that a carefully chosen set of fewer than ten models can preserve most useful behavioral diversity, while nearest-neighbor routers may be accurate yet brittle under small wording changes. Routers based on model descriptions can be less accurate in some settings but much more stable.
Related event: Google DeepMind Proposes Metrics for Model Routing Stability(2 posts)→
More from Research
- LFM2 tokenizer expansion cuts Thai tokens 4× and speeds on-device decoding up to 3.7× — maximelabonne · 2026-07-21
- Argus improves indoor panoramic 3D reconstruction with covisibility and geometry transformers — ducha_aiki · 2026-07-21
- WAIC robots are now hitting commercially useful success rates, says a recap — chris_j_paxton · 2026-07-21
- Unitree launches a remote real-robot benchmark on its own G1 fleet — chris_j_paxton · 2026-07-21
- Adding order metadata makes VLM error detection collapse, new benchmark shows — m_wulfmeier · 2026-07-21
- New prompt template aims to improve spatial reasoning and cut model laziness — legit_api · 2026-07-21