A look at LLM inference routing: why Dynamo does not send requests to the GPU that remembers your prompt

_jaydeepkarale · x · 2026-07-28

The post opens a thread on LLM inference and challenges a common assumption about KV routing: many people think requests should simply be sent to the GPU that already remembers the prompt.

It says Dynamo does not work that way, and hints that the reason is the most interesting part of the system design.

Related event: How NVIDIA Dynamo Prices LLM GPU Routing Instead of Hard KV Rules(14 posts)→

Original post →

More from Infra

Infra channel →