Dynamo scores GPUs by outstanding read work, not request size
toddhooper · x · 2026-07-27
The first term is outstanding read work
The thread says the first term in Dynamo's score is not the request's own size. It is the amount of reading the target GPU still owes:
- Everything it has already agreed to serve but has not finished.
- Your request added on top.
So the question is not "how big are you?" but "how long is the line you're about to join?" That is what makes the system behave like a load balancer rather than a lookup table.
Related event: How NVIDIA Dynamo Prices LLM GPU Routing Instead of Hard KV Rules(14 posts)→
More from Infra
- Together AI adds canary rollouts for zero-downtime model upgrades on dedicated inference — togethercompute · 2026-09-23
- Dedicated Hardware for Running AI Agents at Scale Arrives — cyrilzakka · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23
- Qwen 27B runs 24hr unattended on one RTX5090, builds full Postgres-SpringBoot-React spreadsheet app — anglepoiselife · 2026-09-23
- OpenRoboto Shift launches: decentralized egocentric video data network for robot brains — markjeffrey · 2026-09-23
- Engineer describes designing digital circuits that recycle most of their energy — MikePFrank · 2026-09-23