GKE Pod snapshots cut AI inference cold starts by 89%, loading 70B models in 37s
rseroter · x · 2026-09-22
- Google Cloud introduced GKE Pod snapshots, a feature that saves a workload's running state (including CPU and GPU memory) and restores it on demand instead of cold-starting.
- Cold start reduced by up to 89%: 70B parameter models load in just 37 seconds, 8B models in 15 seconds.
- Pain point: inference servers must download gigabytes of weights into GPU memory (often minutes), and agentic workloads like code execution sandboxes need fast, per-request isolated environments — startup latency hurts UX and prevents rapid autoscaling, forcing overprovisioning.
- Resuming rather than restarting lets infrastructure scale with demand and cuts overprovisioning costs.
More from Infra
- RL infra detail: re-prefill over PipelineRL's KV cache reuse, batch size for GPU utilization — stochasticchasm · 2026-09-22
- Raspberry Pi locks devices to original RAM size, blocking aftermarket memory upgrades — ngxson · 2026-09-22
- fal's H3 Max generates 5 seconds of frontier-quality video in just 3 seconds — gorkem · 2026-09-22
- NVIDIA's EPD Disaggregation Cuts Multimodal TTFT Up to 5x, E2E Latency 7x — dl_weekly · 2026-09-22
- Running MiniMax H3 locally on a 16GB Mac: 8-10s clips in 15-20 minutes — coberholzer · 2026-09-22
- A Wild Async RL Config: 30 Steps x 25K Rollouts Per Step at Parallelism 4 — willcb · 2026-09-22