Baseten tops Coval's voice AI benchmark: STT ~5x faster than OpenAI with lowest WER
baseten · x · 2026-09-09
Coval's voice AI benchmark puts Baseten on the quality-latency Pareto frontier: its STT deployment is roughly 5x faster than OpenAI's while achieving the lowest Word Error Rate. The piece argues production voice AI performance depends on the full inference stack — STT/TTS latency, orchestration, networking, autoscaling, and deployment placement — not just model choice, with multi-model voice agents incurring a 'network tax' of even a few hundred milliseconds.
More from Infra
- Rumor resurfaces: Google may replace Nvidia as TSMC's biggest customer — zephyr_z9 · 2026-09-09
- Tahuna open-sources ephemeral GPU orchestration for ML workloads — Monaim101 · 2026-09-09
- Zuckerberg: Meta already training post-Watermelon models on its 1GW Prometheus cluster — rohanpaul_ai · 2026-09-09
- Developer maps the entire inference + fine-tuning provider landscape with tradeoffs — Present-Jelly9941 · 2026-09-09
- Smart offloading of active requests erases linear attention's memory advantage — samsja19 · 2026-09-09
- Local video generation on Android: ~550-600s per clip on Snapdragon 8 Gen 3 — sgcego · 2026-09-09