Open-source Dynamo deployment guide with benchmarks on a 16×H100 cluster

TheZachMueller · x · 2026-09-20

Junup Park and Jaga Prasanna published an open-source deployment guide for NVIDIA Dynamo on a 2-node, 16×H100 cluster (provided by Lambda). It includes runbooks for standing up K8s/Dynamo from scratch, networking debugging (RoCE, NVLS), model caching, NIXL/Grafana monitoring, and benchmarks on vLLM and SGLang quantifying what P/D disaggregation, KV-aware routing, CPU KV offload, parallelism, and autoscaling actually deliver for production LLM inference.

Original post →

More from Infra

Infra channel →