Google paper: model loading takes 55-70% of cold-start latency for quantized SLMs on serverless CPUs

rohanpaul_ai · x · 2026-09-24

A new Google research paper systematically benchmarks quantized small language models served via llama.cpp on Cloud Run's CPU-only infrastructure.

Key findings:

Scope: 5 models (270M-3.8B), 4/8 GiB memory tiers, and a Q2K through Q80 GGUF quantization sweep. Practical guidance for anyone deploying SLMs on serverless platforms.

Related event: Google Paper: 55-70% of Quantized LLM Cold-Start Latency Is Weight Loading(2 posts)→

Original post →

More from Infra

Infra channel →