NVIDIA Dynamo adds session-aware inference: reuse agent KV cache across vLLM and SGLang
PyTorch · x · 2026-10-09
A PyTorch blog post co-published with NVIDIA details how NVIDIA Dynamo reshapes inference serving for agentic workloads.
The problem: unlike single-turn chat, a coding agent session starts with a large prefill (often tens of thousands of tokens of system prompt, tool definitions and user guidance), resends the full context every turn, and fans out into dozens of model calls plus parallel short- and long-lived subagents. Most wall-clock time is spent waiting on tool calls, with the context sitting resident in the KV cache while nothing generates. Serving efficiency hinges on how much resident context is kept and reused, and how many concurrent sessions fit under sub-second latency targets.
The approach: Dynamo introduces a unified session-level identifier that upgrades request-granularity serving into a program-aware system, enabling:
- Session-aware routing
- Shared KV cache indexing across workers
- Programmatic KV cache movement across vLLM and SGLang engines
The post contrasts chatbot vs. agentic traffic shapes and shows how Dynamo's router, inference engine, and KV-cache manager work end to end.
More from coding & agent
- A Playable 3D Boat Game Embedded in an X Post, Built by Claude Opus — prasenx · 2026-10-09
- edith-1 monitors agent traces, beats Sonnet-5.5 on balanced accuracy at 1/274 the cost — xennygrimmato_ · 2026-10-09
- How edith-1 works: probabilistic filter plus expensive agent judge for flagged runs — xennygrimmato_ · 2026-10-09
- Solari launches agent infrastructure: 8ms browsers, 10x faster than Browserbase — Scobleizer · 2026-10-09
- Monetizing proprietary data through MCP: pay-per-call pricing for agents — ValosSantanos · 2026-10-09
- Splash 1.3.0 cuts local agent first-token latency from 19s to 1s via SSD offloading on M6 Mac — BeidiChen · 2026-10-09