NVIDIA open-sources srt-slurm: YAML orchestration for inference on Slurm clusters

AccBalanced · x · 2026-08-29

SemiAnalysis highlights NVIDIA's open-source srt-slurm, a declarative YAML-based orchestration layer for production-style inference topologies. Running one inference server in a Slurm job is easy; coordinating across multiple containers and nodes is not. srt-slurm coordinates sbatch/srun, GPU placement, networking, readiness checks and cleanup for prefill/decode disaggregation, multiple workers, Dynamo frontends, KV-cache-aware routers, and KV-cache offloading services — useful for quickly deploying representative inference systems for testing or benchmarking on Slurm-based GPU clusters.

Original post →

More from Infra

Infra channel →