Google paper: 55-70% of quantized LLM cold-start latency is just model loading
rohanpaul_ai · x · 2026-09-24
A new Google paper finds that for small quantized LLMs on serverless CPUs, 55-70% of cold-start latency comes from model loading — moving weights into memory, not token generation. Bumping Cloud Run memory from 4GB to 8GB unlocks 2× CPU and nearly halves warm inference time.
Related event: Google Paper: 55-70% of Quantized LLM Cold-Start Latency Is Weight Loading(2 posts)→
More from Infra
- Wasmer's Swift SDK brings sandboxed Python, Node.js and FFmpeg running fully on-device on iOS — jedisct1 · 2026-09-24
- Read-only MCP servers for Proxmox and pfSense: agents hit the API, not screenshots — Mustela__ · 2026-09-24
- From one GPU to millions of users: lessons in LLM inference system design — metalvendetta · 2026-09-24
- Modal Labs in Talks to Raise at $15 Billion Valuation as Inference Rush Heats Up — nmasc_ · 2026-09-24
- Optical computing for AI debated as veteran cites lack of European interest — IgorCarron · 2026-09-24
- glance-vlm speedlab goes open source: MLX 8-bit cuts local camera VLM latency 27.6% — natesiggard · 2026-09-24