Tsinghua's TokenRouter: Token-Level LLM Routing Hits Up to 64.15X Serving Throughput

rohanpaul_ai · x · 2026-10-10

A new Tsinghua paper proposes TokenRouter, a serving system for token-level small/large model routing, reaching up to 64.15X the throughput of existing setups.

Problem: Popular serving frameworks (vLLM, SGLang) run one model per request, so when two models share an answer, every step waits for the slower one.

Approach: TokenRouter gives each model its own server and lets them hand work back and forth — passing a half-written answer between models while preserving the KV cache, and briefly holding requests so each model works on bigger batches.

Results: Across 5 routing methods, throughput rose 2.01–64.15X over the stronger baseline. Paper: arxiv.org/abs/2610.12242.

Related event: Tsinghua's TokenRouter boosts LLM serving throughput 64x(2 posts)→

Original post →

More from Infra

Infra channel →