Nari Labs Open-Sources Qwen3-TTS Engine, Beats ElevenLabs on Accuracy and Price

toebee · hn · 2026-09-15

Nari Labs (makers of Dia) launched Qwen3-TTS/ASR endpoints and open-sourced a custom inference engine for speech models. On Coval benchmarks, their TTS ranks #1 in accuracy (beating ElevenLabs and Cartesia) and is the cheapest endpoint; ASR has the lowest latency and #2 accuracy. Key insight: vLLM/SGLang fit multimodal inference poorly, so they built a specialized engine achieving sub-50ms latency at 10 RPS — outperforming even Alibaba's official endpoints. Next up: diarization, video and world-model inference.

Original post →

More from coding & agent

coding & agent channel →