Snowflake's Semi-Persistence uses CRIU to serve many models on shared GPUs
bookwormengr · x · 2026-09-01
Snowflake AI Research published new work leveraging CRIU for multi-model serving on shared GPUs.
- Problem: serving a single model is easy, but on shared GPUs keeping every model loaded wastes capacity, while swapping only pays off if it's fast enough to stay out of the serving path.
- Approach: Semi-Persistence keeps model weights in pinned CPU memory, balancing GPU utilization and swap speed.
Related event: Snowflake's Semi-Persistence Speeds Up Multi-Model GPU Swapping Up to 19.9x(2 posts)→
More from Infra
- Qwen3.8 Flash hits 415 tok/s on dual DGX Sparks — NVIDIAAI · 2026-09-01
- OpenAI's 'Jalapeno' Chip Revealed: 1500 Tokens/s Throughput — firstadopter · 2026-09-01
- TensorSharp vs llama.cpp: Qwen 3.8 Flash Next Benchmarks — fuzhongkai · 2026-09-01
- Tencent Hunyuan AngelSlim: Compressing Hy4 Model to 214GB with Heterogeneous Inference — 腾讯混元 · 2026-09-01
- Samsung shifts to 8-layer HBM4E for Nvidia with ~20% higher speed spec — 创业邦 · 2026-09-01
- Why did increasing context size increase speed in Llama.cpp? — satnl · 2026-09-01