Dynamo turns routing into one cache-block cost per GPU

Abhishekcur · x · 2026-07-27

Dynamo collapses queueing, space, and cache into one score

This thread's core trick is to stop treating queueing, occupied space, and cache reuse as separate decisions.

In other words, there is no "check memory first, then load" logic. There is just one unified cost function, and the rest is deciding what belongs in it.

Related event: How NVIDIA Dynamo Prices LLM GPU Routing Instead of Hard KV Rules(14 posts)→

Original post →

More from Infra

Infra channel →