Dynamo turns routing into one cache-block cost per GPU
Abhishekcur · x · 2026-07-27
Dynamo collapses queueing, space, and cache into one score
This thread's core trick is to stop treating queueing, occupied space, and cache reuse as separate decisions.
- All three are measured in the same unit: cache blocks.
- Once everything is normalized, each GPU gets a single number.
- The router simply picks the smallest score.
In other words, there is no "check memory first, then load" logic. There is just one unified cost function, and the rest is deciding what belongs in it.
Related event: How NVIDIA Dynamo Prices LLM GPU Routing Instead of Hard KV Rules(14 posts)→
More from Infra
- Dedicated Hardware for Running AI Agents at Scale Arrives — cyrilzakka · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23
- Qwen 27B runs 24hr unattended on one RTX5090, builds full Postgres-SpringBoot-React spreadsheet app — anglepoiselife · 2026-09-23
- OpenRoboto Shift launches: decentralized egocentric video data network for robot brains — markjeffrey · 2026-09-23
- Engineer describes designing digital circuits that recycle most of their energy — MikePFrank · 2026-09-23
- Cloudflare CTO Dane Knecht makes TIME's 2026 executives list as AI crawlers hit 52% of traffic — dinasaur_404 · 2026-09-23