Cloud Run's hidden 60% scaling dials go GA: CPU and concurrency targets now customizable
rseroter · x · 2026-10-07
Google Cloud Run services have long auto-scaled on a fixed, invisible 60% CPU utilization target plus a 60% concurrency target. Custom scaling controls introduced in preview on April 16, 2026 went GA on September 29.
Key points:
- CPU target is now adjustable from 0.10–0.90 and concurrency target from 0.10–0.95, both defaulting to 60%. You can enable either one or neither (but not turn both off entirely).
- The concurrency target is a fraction of max concurrent requests, not a replacement for it. Whichever of CPU, concurrency, or adaptive concurrency demands more instances wins.
- Lower targets scale earlier at higher cost; higher targets create a wider dead zone with step changes.
- Billing: request-based billing counts CPU only during requests; instance-based counts the full instance lifetime, except scale-to/from-zero which remains request-only.
- Applies to services only, not jobs or worker pools.
More from Infra
- SpaceX reportedly seeks $40B to fund Nvidia AI chip purchases, Apollo leading — Polymarket · 2026-10-07
- Dev warns OpenRouter share, cache hit rate, latency stats are easily gamed for marketing — charles_irl · 2026-10-07
- Dan Ives: investors underestimate a $4T tech spending wave; top picks Nvidia, Microsoft, Palantir — matt_slotnick · 2026-10-07
- 27B Model at 256k Context, 110+ tok/s on a Single RTX 5090 via focus-llama — Ok-Shower7286 · 2026-10-07
- Running Qwen3.8 27B on 16GB VRAM: IQ3 quant hits 131k context at 9-30 tps — randomgenericbot · 2026-10-07
- "Pre-training is now just a warmup": post-training RL and inference take over compute — joeddav · 2026-10-07