OpenAI Post-Mortem: Telemetry Service Overwhelmed K8s Control Plane Causing Global Outage
techNmak · x · 2026-08-26
- Overview: On December 11, 2024, OpenAI experienced a significant downtime affecting all services for over 4 hours.
- Root Cause: A new telemetry service deployment intended to improve observability accidentally caused thousands of nodes to hammer the Kubernetes API servers, overwhelming the control plane.
- Cascading Failures: DNS caching delayed the visibility of the failure. Critically, the broken control plane was also required for recovery efforts, complicating remediation.
- Recovery: Substantial recovery for ChatGPT started at 5:45 PM PST and for API at 5:36 PM PST, with full resolution across all models by 7:38 PM PST.
- Key Takeaways: Telemetry components can act as internal DoS sources; control plane health must be decoupled from recovery paths; DNS caching can mask early failure indicators.
More from Infra
- Two vLLM recipes for Blackwell: NVFP4 KV cache buys 262K context and more streams — SeanHighness · 2026-08-27
- Dev forks Nvidia drivers to enable PCIe P2P on GeForce for SlimServe — QuixiAI · 2026-08-27
- View: Single dev with 200 B300s could beat Alibaba's post-training team — kalomaze · 2026-08-27
- GLM 5.3 Flash Benchmark: Hits 881 tok/s on Dual DGX — teortaxesTex · 2026-08-27
- Rakyll advises: If you have any CPU nodes, hold on to them — rakyll · 2026-08-27
- OpenRouter Serves 200B Total Tokens, Adds Qwen 3.5 35B — gajesh · 2026-08-27