Dynamo Architecture: Pricing Memory to Optimize GPU Routing
Abhishekcur · x · 2026-07-28
The author provides a detailed breakdown of the clever GPU routing mechanism in the Dynamo architecture, which transforms memory from a rigid rule into a dynamic price.
- Price over Rule: Because the answer is a number, the system can loosen limits, allowing tasks to favor cheaper GPUs. This naturally spreads heat without ruining the ordering.
- Lock-free Checking: Since calculating the price changes nothing and reserves no resources, the router can freely check every GPU in the cluster, committing only at the very end.
- Coin-flip Tie-breaking: When a real tie occurs, it’s settled by a coin flip, preventing a fresh cluster from quietly piling tasks onto whichever GPU gets checked first.
- Dynamic Discount Shrinking: The smartest part is that as a GPU's queue gets longer, the discount for its local cache shrinks, while discounts for normal memory or disk cache remain whole.
Related event: How NVIDIA Dynamo Prices LLM GPU Routing Instead of Hard KV Rules(14 posts)→
More from Infra
- Dedicated Hardware for Running AI Agents at Scale Arrives — cyrilzakka · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23
- Qwen 27B runs 24hr unattended on one RTX5090, builds full Postgres-SpringBoot-React spreadsheet app — anglepoiselife · 2026-09-23
- OpenRoboto Shift launches: decentralized egocentric video data network for robot brains — markjeffrey · 2026-09-23
- Engineer describes designing digital circuits that recycle most of their energy — MikePFrank · 2026-09-23
- Cloudflare CTO Dane Knecht makes TIME's 2026 executives list as AI crawlers hit 52% of traffic — dinasaur_404 · 2026-09-23