Infinigence's PDD Architecture Halves First-Token Latency, Cuts Costs by 37%

量子位 · wechat · 2026-07-30

At WAIC 2026, Infinigence unveiled its cross-cluster heterogeneous inference architecture, PDD. It connects local homogeneous data centers via low-cost WANs and evolves the traditional Prefill-Decode (PD) separation into a three-tier "P-RLD-MD" structure. This architecture overcomes KVCache transmission latency bottlenecks in Ethernet environments, reducing first-token latency by 51.5% and single-token costs by 37.5% in tests.

The team detailed three core insights driving the design: extreme hardware specialization, a 10x bandwidth reduction via prefix caching (DRC), and non-uniform latency caused by a tiny fraction of low-hit-rate requests in agent workloads. To solve this, PDD introduces a "relay" mechanism using local RLD instances to mask cross-cluster network delays. By transferring only Token IDs and recomputing locally, it converts unpredictable network transfers into deterministic computations, ultimately achieving a 37.5% improvement in cost-efficiency.

Original post →

More from Infra

Infra channel →