Optimizing Qwen3-ASR Latency to 70ms to Beat Deepgram

Comprehensive_Quit67 · reddit · 2026-08-25

The author optimized the Qwen3-ASR 1.7B pipeline, reducing streaming latency from 400ms (via Baseten/official) to 70ms. This allows it to surpass Deepgram in speed while maintaining superior accuracy in WER and multilingual benchmarks, without changing model weights.

Original post →

More from Infra

Infra channel →