OpenAI details scaling to 70M requests/sec over 500PB across nearly 40 regions
AxSaucedo · x · 2026-09-16
OpenAI published a two-part infrastructure series explaining how its serving stack supports ChatGPT at scale: 70 million requests per second, over 500 petabytes, across nearly 40 regions, serving more than a billion people per week. The posts detail the engineering behind its globally distributed inference infrastructure.
More from Infra
- Transformers models now run natively in vLLM with no port required — pcuenq · 2026-09-16
- TSMC builds 20 fabs yet can't meet AI demand as labor shortage slows expansion — emmanuelvivier · 2026-09-16
- SemiAnalysis: Nvidia Vera Rubin NVL72 Delivers Up to 30x Higher Throughput per MW Than Blackwell for Agentic Inference — emmanuelvivier · 2026-09-16
- Cacheon Miners Push MiniMax M3 to 2,337.9 tok/s, +32.2% Over SGLang — Modeled as ~61% More Profit per Chip-Hour — JosephJacks_ · 2026-09-16
- Paying three AI vendors to parse our own docs: a Reddit quest for one self-hosted stack — Sad-Razzmatazz-7657 · 2026-09-16
- Nadella: a 400-500MW data center grew a rural town's tax revenue 12x — rohanpaul_ai · 2026-09-16