Why Agentic Inference Needs P/D Disaggregation: Decoding 300K-Token Codebase Latency
abhijithneil · x · 2026-10-01
ekzhang1 (with Horace and the Thinky team) shared an analysis of why P/D disaggregation helps, and jagaprasanna added an engineering perspective.
Key takeaways
- The real benefit of P/D disagg is improving mean per-token latency in steady state, not fixing single-token tail latencies; it helps most for prefill-heavy workloads, not ultra-low-latency use cases.
- Agentic inference (PR reviews, codebase analysis, long-running coding agents) is a very different workload from standard chat: think a 300K-token codebase context plus 50K output tokens, which is extremely prefill-hungry.
- On aggregated workers with chunked prefill, the same workers handle heavy prefills and ongoing decode simultaneously, spiking ITL (inter-token latency).
- Standard chatbots have smaller contexts, fewer repeated tool calls, and balanced loads, so aggregation can make more sense there — explaining why P/D disaggregation matters especially for agentic and coding workloads.
More from coding & agent
- NVIDIA's Mid-Harness Scales Actions at the Model-Harness Boundary, Lifting TerminalBench Pass@1 to 68.03% — nvidia · 2026-10-01
- Amazon's SMART Self-Evolving Multi-Agent System Tops All 15 Subtitle Arena Directions, Cuts Penalty 6.9% — amazon · 2026-10-01
- Gary Bernhardt hits all-time low faith in AI agents: they "fix" tests by deleting them — sidjustice_ · 2026-10-01
- 'Agents are the software now' — developer urge to dive in — PurzBeats · 2026-10-01
- Hybris MCP Server lets AI assistants manage SAP Commerce Cloud instances — modelcontextprotocol · 2026-10-01
- cua-speedrun: CMU benchmark shows 4.4x speed gap between equal-scoring computer-use agents — arankomatsuzaki · 2026-10-01