SaveRouter: Sparse Supervision Cuts LLM Router Training Cost, Break-Even Volume Down 9.5x
SinapisAI · hf · 2026-09-30
The SaveRouter paper tackles the economics of LLM routing: learning a router typically requires executing multiple candidate models on historical queries to collect quality feedback — an upfront supervision cost that existing serving-efficiency work overlooks. Routing quality also tends to saturate before all query–model feedback is collected.
SaveRouter is a sparse-supervision framework that selectively acquires informative model feedback, shares capability information across related queries, and retains query-level refinement. Across four routing benchmarks:
- Uses only 33–41% of available training feedback while matching or beating routing quality
- Reduces break-even deployment volume by 1.9–9.5x versus the fastest conventional router
- More supervision isn't always economically preferable: the level minimizing serving cost can differ from the one achieving earliest payback
Code is open-sourced at LAMDA-Model-Reuse/SaveRouter.
More from Infra
- Dual R9700 local inference: how much does PCIe Gen4 actually cost vs Gen5? — IngwiePhoenix · 2026-09-30
- SwitchSD, speculative decoding that switches between neural drafting and context copying, accepted to NeurIPS — arankomatsuzaki · 2026-09-30
- Rebellions' 2048 TFLOPS NPU lands major Japanese AI datacenter deal with ai& — DavidBennett__ · 2026-09-30
- Solo Rust + Vulkan training backend now passes 14 architectures at 2e-7 parity, with full LoRA/PEFT lifecycle — PhysicsDisastrous462 · 2026-09-30
- Micron guided $50B quarterly revenue; independent model says $56.2B if price hikes pass through — tengyanAI · 2026-09-30
- DeepSeek ships open-source TileLang tools for Huawei Ascend chips, countering CUDA — The Decoder · 2026-09-30