Serverless AI: why usage-based billing can cost more for static workloads
DavidLinthicum · x · 2026-09-08
Cloud architect David Linthicum lays out the case for and against Serverless AI.
Key points:
- Serverless has been the default architecture for resilience over the past decade, and managed AI APIs — Amazon Bedrock, Azure OpenAI, Google Vertex AI — now let teams skip GPU provisioning, model deployment, and capacity planning entirely
- The trap many architects fall into: usage-based billing doesn't always mean lower costs. For AI applications with a static, steady resource footprint over time, serverless typically costs more than self-managed infrastructure
- Takeaway: match the deployment model to workload shape (steady vs. spiky) rather than defaulting to serverless
More from Infra
- 8GB VRAM runs Flux and Wan 2.2 locally fine: three wrong settings, not the GPU, were the bottleneck — leonbuilds · 2026-09-08
- China effectively leads the humanoid robot supply chain, and Optimus relies on it — JOBhakdi · 2026-09-08
- Benchmarked: Apple Core AI vs MLX for on-device LLM speed on iPhone and Mac — HankYeomans · 2026-09-08
- Tiered KV-cache offloading for self-hosted LLM inference: GPU to RAM to NVMe to S3 — Responsible-You9024 · 2026-09-08
- D-Wave Finalizes Agreement with US Commerce Dept for Up to $100M in CHIPS Act Funding — ceciletamura · 2026-09-08
- Export Controls Working? H200 Sells for 280 and B300 for 450 Overseas — teortaxesTex · 2026-09-08