Google paper: 55-70% of quantized LLM cold-start latency is just model loading

rohanpaul_ai · x · 2026-09-24

A new Google paper finds that for small quantized LLMs on serverless CPUs, 55-70% of cold-start latency comes from model loading — moving weights into memory, not token generation. Bumping Cloud Run memory from 4GB to 8GB unlocks 2× CPU and nearly halves warm inference time.

Related event: Google Paper: 55-70% of Quantized LLM Cold-Start Latency Is Weight Loading(2 posts)→

Original post →

More from Infra

Infra channel →