Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter
Ruoyu Qin, Weiran He, Yaoyu Wang, Zheming Li, Xinran Xu, Yongwei Wu, Weimin Zheng, Mingxing Zhang
cs.DC
2026-04-16
PrfaaS offloads long prefill to remote compute clusters over Ethernet; on a 1T hybrid model, +54% throughput and -64% P90 TTFT vs homogeneous PD, ~15% more at equal cost.
Prefill-decode (PD) disaggregation is now the default for large-scale LLM serving: prefill burns compute, decode burns memory bandwidth. Moonshot's Mooncake made KVCache a first-class resource. The obvious next step is heterogeneous silicon, compute-dense chips for prefill and bandwidth-optimized chips for decode. NVIDIA's Rubin CPX and Groq's LPU are already built that way.
The binding constraint is KVCache transport. Under dense attention, KV grows linearly with length. A MiniMax-M2.5 instance on 8×H200 emits about 60 Gbps of KV at 32K tokens, which commodity inter-datacenter Ethernet cannot carry, so both phases stay inside one RDMA fabric. Specialized chips are usually bought and sited by type. Forcing them into one cluster also freezes the prefill-to-decode hardware ratio, and when traffic mix shifts one side sits idle.
Hybrid attention changes the arithmetic. A few full-attention layers interleaved with linear or sliding-window layers can cut KV volume by about an order of magnitude, so cross-datacenter transfer becomes plausible. Arrivals are still bursty, lengths are skewed, prefix caches are uneven, and links jitter. Shipping every prefill still congests the pipe. Smaller KV opens the door; the system still has to pick which requests are worth sending.
PrfaaS (Prefill-as-a-Service) offloads only long, uncached prefills to a standalone compute-dense cluster and ships the resulting KVCache over Ethernet to a local PD cluster for decode. Short requests stay on the local prefill path. Heterogeneous accelerators no longer need to share a low-latency RDMA island.
Length-threshold routing is the first filter. Let l be the uncached increment and t the threshold: l > t goes remote. Short prefills are usually memory- or communication-bound, so they waste compute-dense chips and, per unit time, emit more KV, which fills the egress link sooner.
The scheduler runs on two timescales. In the short run it watches PrfaaS egress utilization and queue depth, and raises t when the link nears its ceiling. On prefix hits it also looks at where the cache lives: if bandwidth is tight, each cluster uses only its own prefix; if bandwidth is slack, the longer prefix can move across clusters to skip redundant compute. In the long run it retunes the local PD split between prefill and decode roles and recomputes t.
The cache pool splits two kinds of state. Linear-attention or SWA recurrent states are request-level and fixed-size; they reuse only on an exact-length match. Full-attention KV is block-level and supports partial prefix hits. The two live in separate groups over a shared block pool, tagged as reusable prefix blocks or transfer blocks that are dropped after the cross-cluster send. Layer-wise prefill pipelining overlaps compute with send; multi-connection TCP fills Ethernet; loss and retransmission signals feed back into the scheduler.
A throughput model treats PrfaaS prefill, local PD-P, and local PD-D as three stages and takes the slowest after adjusting for the offload fraction p. The optimum saturates both prefill paths at once and matches their sum to decode. A grid search over t and Np/Nd finds that point.
Open models quantify the bandwidth gap. At 32K tokens on 8×H200 with SGLang v0.5.9, MiMo-V2-Flash emits 4.66 Gbps of KV versus 59.93 Gbps for MiniMax-M2.5, about 13×. Qwen3.5-397B sits at 8.25 Gbps versus 33.35 Gbps for dense Qwen3-235B, about 4×. For Ring-2.5-1T, MLA compresses roughly 4.5× relative to GQA and a 7:1 hybrid ratio adds about 8×, for 36× less KV memory. On 512 H200 GPUs with 32K average uncached input, MiniMax-M2.5 wants 3.8 Tbps of egress and Qwen3 wants 2.1 Tbps; Ring-2.5-1T needs about 170 Gbps, and routing 128K-class requests to PrfaaS can push that under 100 Gbps.
The case study uses an internal 1T hybrid model with KDA:MLA at 3:1, the same layout as Kimi Linear. Two clusters share 100 Gbps of VPC: 32 H200 GPUs for PrfaaS, 64 H20 GPUs for local PD (800 Gbps RDMA per node), against a 96-H20 homogeneous PD baseline. Input lengths follow a truncated log-normal (µ=9.90, σ=1.00, [128, 128K]) with mean 27K; output is fixed at 1024 tokens; SLO is 40 tokens/s, excluding speculative decoding. Throughput and bandwidth numbers come from feeding measured profiles into the §3.4 model.
The operating point is t=19.4K, with 49.6% of requests offloaded and mean offloaded length 44K. Local PD runs 3 prefill instances and 5 decode instances (8 GPUs each).
| Setup | Λmax (req/s) | vs homogeneous | Mean TTFT | P90 TTFT |
| PrfaaS-PD (4/3/5) | 3.24 | 1.54× | 2.22 s | 3.51 s |
| Homogeneous PD (9/3) | 2.11 | 1.00× | 4.44 s | 9.73 s |
| Naive heterogeneous (4/-/8) | 2.45 | 1.16× | 1.74 s | 3.51 s |
Versus homogeneous PD, throughput is 54% higher, P90 TTFT 64% lower, mean TTFT 50% lower. Naive heterogeneous serving puts all prefill on H200 and all decode on H20 with no length routing, and only reaches 1.16×, about 25% below the scheduled design (3.24 vs 2.45). Average PrfaaS egress is 13 Gbps, 13% of the 100 Gbps link. At equal cost the paper reports about 15% more throughput. H200 and H20 are a stand-in pair.
Once hybrid attention brings per-instance KV throughput down to single-digit Gbps, heterogeneous PD no longer has to co-locate both chip types behind one RDMA fabric. Compute clusters and bandwidth clusters can scale in different buildings, and long requests stop competing with short ones for the same prefill slots.
Dense GQA models do not get this path. MiniMax-M2.5 still sits near 60 Gbps per instance at 32K, which is Tbps-scale egress at 512 GPUs. The design assumes a hybrid stack (KDA, SWA, GDN, and similar) plus selective offload. For a team already running Mooncake-style PD and moving the next model to hybrid attention, this reads as a capacity plan.
Part of the 54% is simply H200 being faster. The scheduling increment is cleaner: 32% over naive heterogeneous. Equal-cost gain is only about 15%, so remote prefill plus a chip swap is an incremental cost win.
Every number in §4 comes from profiles plus an analytical model, not end-to-end measurements on two clusters under live bursty traffic. Congestion and prefix-cache imbalance are central in the design sections and absent as variance in the result table. Output length is pinned at 1024 and SLO at 40 tokens/s; agentic incremental prefills would move the optimal t.
The 54% figure does not separate silicon from systems. The 15% equal-cost claim is approximate, with no public unit-price assumption. Naive heterogeneous already matches PrfaaS-PD on P90 TTFT at 3.51 s; scheduling mainly buys throughput and mean TTFT. Layer-wise pipelining, multi-connection TCP, and cross-cluster cache moves have no ablations. The 1T model and traffic mix are internal, so outsiders can only cross-check Table 3's open Φkv numbers.
The paper notes that the PrfaaS cluster is still compute-bound with ample bandwidth headroom, so larger deployments can add more prefill GPUs. If prefill chips get much faster, the 13 Gbps headroom disappears first, and bandwidth-aware scheduling becomes a hard constraint.