Baseten tops Coval's voice AI benchmark: STT ~5x faster than OpenAI with lowest WER

baseten · x · 2026-09-09

Coval's voice AI benchmark puts Baseten on the quality-latency Pareto frontier: its STT deployment is roughly 5x faster than OpenAI's while achieving the lowest Word Error Rate. The piece argues production voice AI performance depends on the full inference stack — STT/TTS latency, orchestration, networking, autoscaling, and deployment placement — not just model choice, with multi-model voice agents incurring a 'network tax' of even a few hundred milliseconds.

Original post →

More from Infra

Infra channel →