TokenRouter Serving System Boosts Token-Level LLM Routing Throughput up to 64x
nics-efc · hf · 2026-10-09
Researchers from THU's nics-efc lab released TokenRouter, a serving system for token-level LLM routing. It follows request-centric programming with model-centric execution: developers describe routing logic per request while the runtime spins up a subserver per LLM with asynchronous dispatch and a delayed-batching scheduler tuned via a mathematical throughput model. Across routing algorithms, workloads, and model pairs, it achieves 2.01–64.15x higher decoding throughput than existing systems. Code is open-sourced.
More from Infra
- OpenAI revenue definition gap sparks AI hardware selloff as TSMC posts +55% YoY — tengyanAI · 2026-10-09
- LoRA over GGUF: fine-tune Qwen3.8-Flash-Next in 40 GiB VRAM without CPU offloading — woct0rdho · 2026-10-09
- Nunchux inference engine joins AMD's AI Inference Engines & Services ecosystem — junyanz89 · 2026-10-09
- tinygrad runs its GitHub Actions CI on 4 new tinybox machines — AIFlow_ML · 2026-10-09
- OpenAI's 10,000-agent, 130B-token run pushed slime v0.4.0 to rethink RL infrastructure scale — teortaxesTex · 2026-10-09
- SparseDecoding: Decoding-Aware Pruning Yields up to 1.48x Faster LLM Inference — encodelab · 2026-10-09