Tsinghua's TokenRouter: Token-Level LLM Routing Hits Up to 64.15X Serving Throughput
rohanpaul_ai · x · 2026-10-10
A new Tsinghua paper proposes TokenRouter, a serving system for token-level small/large model routing, reaching up to 64.15X the throughput of existing setups.
Problem: Popular serving frameworks (vLLM, SGLang) run one model per request, so when two models share an answer, every step waits for the slower one.
Approach: TokenRouter gives each model its own server and lets them hand work back and forth — passing a half-written answer between models while preserving the KV cache, and briefly holding requests so each model works on bigger batches.
Results: Across 5 routing methods, throughput rose 2.01–64.15X over the stronger baseline. Paper: arxiv.org/abs/2610.12242.
Related event: Tsinghua's TokenRouter boosts LLM serving throughput 64x(2 posts)→
More from Infra
- DuckDB v2.0 CLI agent mode cuts agent-read tokens by 59% on TPC-H benchmarks — josh_wills · 2026-10-10
- Datology releases Zephon, a deterministic on-the-fly dataloader born from MosaicML Streaming's legacy — josh_wills · 2026-10-10
- Meta Muse Auto-Routes to OpenRouter Free Models for Zero-Cost Long Tasks — sven_ai · 2026-10-10
- DIY hybrid GPU/CPU/SSD rig cuts DeepSeek TTFT from 75s to 8.9s at 16K prefill — HankYeomans · 2026-10-10
- vLLM thread (4/5): locality-domain MoE sharding speeds up decode 1.2x — vllm_project · 2026-10-10
- vLLM lands NVIDIA Vera Rubin support, hitting 7.8x GB200 throughput on MiniMax M3 — vllm_project · 2026-10-10