Infinigence's PDD Architecture Halves First-Token Latency, Cuts Costs by 37%
量子位 · wechat · 2026-07-30
At WAIC 2026, Infinigence unveiled its cross-cluster heterogeneous inference architecture, PDD. It connects local homogeneous data centers via low-cost WANs and evolves the traditional Prefill-Decode (PD) separation into a three-tier "P-RLD-MD" structure. This architecture overcomes KVCache transmission latency bottlenecks in Ethernet environments, reducing first-token latency by 51.5% and single-token costs by 37.5% in tests.
The team detailed three core insights driving the design: extreme hardware specialization, a 10x bandwidth reduction via prefix caching (DRC), and non-uniform latency caused by a tiny fraction of low-hit-rate requests in agent workloads. To solve this, PDD introduces a "relay" mechanism using local RLD instances to mask cross-cluster network delays. By transferring only Token IDs and recomputing locally, it converts unpredictable network transfers into deterministic computations, ultimately achieving a 37.5% improvement in cost-efficiency.
More from Infra
- Run Gemma 4 Locally with 16GB RAM: A Zero-Cost Fully Offline Setup Guide — FinanceYF5 · 2026-07-30
- 4090+5060 Ti Hybrid Inference: Runs 122B Model at 37 t/s — Dry_Long3157 · 2026-07-30
- OpenAI to Consume 40% of Global DRAM: The Rise of the Metered Intelligence Complex — Shimano-No-Kyoken · 2026-07-30
- AI Giants Accused of Hiding $1.65 Trillion in Off-Balance-Sheet Debt, Echoing Enron — marigo · 2026-07-30
- Ollama Partners with Intel to Enable Local LLM Inference on Core Ultra Series 3 — ollama · 2026-07-30
- New S3 Client Delivers 20x Throughput Increase on a Single Core — mgill25 · 2026-07-30