Reusing connections cut LLM pre-call budget check p99 from 1278ms to 414ms
Tough_Stretch_4045 · reddit · 2026-09-23
A solo developer measured the budget-check service (Redis reads before every LLM call, in us-east-1) from India: reusing one connection gave p50 317ms / p99 414ms over 200 prod calls, vs p50 446ms / p99 1278ms with a new connection each time. The gap is almost entirely TCP+TLS handshake (156ms + 315ms cold; 253ms warm floor from distance). Surprise finding: cf-ray shows traffic routed via Cloudflare's Marseille edge, not Mumbai/Chennai, adding uncontrollable routing latency. The client gives up at 2s and lets the call through, so worst case is an unchecked call. He asks others running gateways/guardrails what latency they accept.
More from Infra
- DarkbloomAI one month paid on OpenRouter: 4B to 20B+ tokens/day — gajesh · 2026-09-23
- Programmable Si photonic circuit hits 29 fW static power per pi phase shift — jwt0625 · 2026-09-23
- Clean pre-dicing photonic wafer shows low-power InGaAsP-on-silicon phase modulators — jwt0625 · 2026-09-23
- Gas turbine orders booked to 2030, prices up 195% as AI power gap widens — FinanceYF5 · 2026-09-23
- US datacenters need 18GW in 2026 but face ~5GW shortfall after fixes — FinanceYF5 · 2026-09-23
- Morgan Stanley: US datacenter power gap equals six New York Cities — FinanceYF5 · 2026-09-23