Snowflake's Semi-Persistence Cuts Multi-Model GPU Sleep/Wake Cycles by Up to 19.9x
StasBekman · x · 2026-09-01
Snowflake AI Research's Semi-Persistence tackles multi-model serving on shared GPUs: keeping every model loaded wastes capacity, but swapping only helps if it stays out of the serving path.
The approach keeps model weights in a pinned CPU-memory pool and treats the GPU copy as temporary; when demand returns, weights stream back over PCIe and NVLink in parallel. Across models from 2B to 397B parameters, internal benchmarks showed 5.6x–19.9x faster sleep/wake cycles vs. the baseline vLLM path, with sub-second swaps for single-GPU models.
Related event: Snowflake's Semi-Persistence Speeds Up Multi-Model GPU Swapping Up to 19.9x(2 posts)→
More from Infra
- Ollama Details Transparent Pricing: No Hidden Fees, Team Plan Live — ollama · 2026-09-01
- Ollama Switches to Transparent Per-Token Pricing with Monthly Credit Pools — ollama · 2026-09-01
- Omarchy achieves first Linux TouchID crack on T1 MacBooks — DanWahlin · 2026-09-01
- Question: Have data providers started training their own models? — xeophon · 2026-09-01
- Stop leaving your AI Agent running 24/7: Power management guide for developers — Rhishi99 · 2026-09-01
- Huge price gaps in Token resources: self-deployed GLM and DeepSeek available at up to 80% off — lipeng0820 · 2026-09-01