SGLang's Weight Cache Daemon Cuts 1T Model Restart Time from 8.8 min to 32 sec

xiaosun86 · x · 2026-08-22

SGLang, in collaboration with Ant Group and Alibaba, introduced the Weight Cache Daemon to drastically reduce model reload times. By keeping post-quantized weights persistent in GPU memory and using CUDA IPC zero-copy mapping, the new system achieves:

Key Features:

This is the first phase of their Fast Engine Recovery Framework, targeting <10s cold restarts and <1s warm standby switches for production LLM serving.

Original post →

More from Infra

Infra channel →