SaveRouter: Sparse Supervision Cuts LLM Router Training Cost, Break-Even Volume Down 9.5x

SinapisAI · hf · 2026-09-30

The SaveRouter paper tackles the economics of LLM routing: learning a router typically requires executing multiple candidate models on historical queries to collect quality feedback — an upfront supervision cost that existing serving-efficiency work overlooks. Routing quality also tends to saturate before all query–model feedback is collected.

SaveRouter is a sparse-supervision framework that selectively acquires informative model feedback, shares capability information across related queries, and retains query-level refinement. Across four routing benchmarks:

Code is open-sourced at LAMDA-Model-Reuse/SaveRouter.

Original post →

More from Infra

Infra channel →