Qwen3-TTS on one H100: sub-50ms p95 latency at ~$2 per 1M characters
bibryam · x · 2026-08-31
Nari Labs released their own serving implementation and benchmark for Qwen3-TTS 1.7B CustomVoice: on a single NVIDIA H100 SXM it hits 10 RPS with sub-50 ms p95 time-to-first-audio and zero underruns during real-time playback.
- Compared against vLLM-Omni, SGLang-Omni, VoxServe, and M under 5-minute Poisson open-loop traffic, theirs is the only implementation achieving sub-50 ms p95 TTFA, staying under 100 ms even at 20 RPS.
- Cost: 630 chars/sec at 10 RPS; at $4.29/hr for the H100 that's roughly $2 per 1M characters — versus $100/1M for ElevenLabs V3 and $49/1M for Cartesia Sonic 3.5 at higher latency.
- Implementation and benchmark are open source; they also define the four requirements of "real-time" TTS: low audible TTFA, zero underruns, capacity as RPS scales, non-malformed output.
More from Infra
- Dual-GPU AI Workstation: R9700 for H3 + 3060 for Qwen Benchmark — Master-Client6682 · 2026-08-31
- Gavin Baker: Why AI Demand Is Outrunning Compute Supply — a16z Podcast · 2026-08-31
- Samsung allocates 70% of memory capacity through 2031 to LTAs, plans conversion — zephyr_z9 · 2026-08-31
- OpenAI Rivals Buy Tens of Thousands of Mac Minis for Agent Training — The Decoder · 2026-08-31
- Running Qwen3.8-Flash-Next on a 96GB Mac Studio: A Deep Dive — Mxmtm · 2026-08-31
- Seeking advice for local dev setup on dual RTX 6000s — alexp702 · 2026-08-31