Real-time Voice Dev Seeks Low P99 Latency EU-Hosted LLM Providers
mogottsch · reddit · 2026-08-21
A real-time voice agent developer is seeking recommendations for LLM providers that meet the following criteria:
- Region: Must be hosted in the EU for data residency compliance.
- Performance: Requires extremely low P99/P95 Time To First Token (TTFT). Current tested providers show good median latency but suffer from severe tail latency spikes.
Tested Providers:
- Vertex AI / Gemini
- Melious / GLM
- Azure OpenAI
- Cerebras
- Scaleway
- Anthropic
Required Features: Streaming and Tool Calling support.
More from Infra
- Can Chinese LLMs Offer 100T Tokens for Free? Cost Advantages Spark Speculation — teortaxesTex · 2026-08-21
- Data Center Moratorium Wave: Multiple Governors and Hundreds of Local Actions, Tracking Dashboard Launched — kevinsxu · 2026-08-21
- Anaconda Releases Best AI Development Tools Guide for 2026 — anacondainc · 2026-08-21
- Localhost-only MCP server used cross-machine via reverse encrypted tunnel, zero exposed ports — XVX109 · 2026-08-21
- Study: LLMs cost 1431x more, embeddings win on classification — vboykis · 2026-08-21
- SK Hynix aims to solve AI bottlenecks with light-based communication — pstAsiatech · 2026-08-21