NVIDIA Dynamo adds session-aware inference: reuse agent KV cache across vLLM and SGLang

PyTorch · x · 2026-10-09

A PyTorch blog post co-published with NVIDIA details how NVIDIA Dynamo reshapes inference serving for agentic workloads.

The problem: unlike single-turn chat, a coding agent session starts with a large prefill (often tens of thousands of tokens of system prompt, tool definitions and user guidance), resends the full context every turn, and fans out into dozens of model calls plus parallel short- and long-lived subagents. Most wall-clock time is spent waiting on tool calls, with the context sitting resident in the KV cache while nothing generates. Serving efficiency hinges on how much resident context is kept and reused, and how many concurrent sessions fit under sub-second latency targets.

The approach: Dynamo introduces a unified session-level identifier that upgrades request-granularity serving into a program-aware system, enabling:

The post contrasts chatbot vs. agentic traffic shapes and shows how Dynamo's router, inference engine, and KV-cache manager work end to end.

Original post →

More from coding & agent

coding & agent channel →