Qwen3-TTS on one H100: sub-50ms p95 latency at ~$2 per 1M characters

bibryam · x · 2026-08-31

Nari Labs released their own serving implementation and benchmark for Qwen3-TTS 1.7B CustomVoice: on a single NVIDIA H100 SXM it hits 10 RPS with sub-50 ms p95 time-to-first-audio and zero underruns during real-time playback.

Original post →

More from Infra

Infra channel →