NVIDIA Dynamo Snapshot cuts LLM serving cold-start times by ~10x

Nasereliver · x · 2026-10-01

A SysConf 2026 talk announcement: Rasheedat Atinuke Jamiu will present her experiments on fixing the LLM cold-start problem with NVIDIA Dynamo Snapshot.

Scaling up LLM serving is slow because new GPU workers must initialize the runtime, load model weights, prepare kernels, and allocate memory. Dynamo Snapshot checkpoints an initialized worker and restores new workers from it — using CRIU for host state and cuda-checkpoint for GPU state.

Her experimental results show roughly a 10x drop in startup times.

Original post →

More from Infra

Infra channel →