NVIDIA maps the streaming multi-speaker ASR design space across four architectures, Interspeech 2026

alexcovo_eth · x · 2026-09-12

A NVIDIA research team (Taejin Park, Ivan Medennikov, Boris Ginsburg, et al.) published a systematic study of streaming multi-speaker ASR, accepted to Interspeech 2026.

The paper categorizes the field into four architectural strategies based on how diarization and ASR are integrated: with/without multiple model instances and with/without fine-tuning. Using a shared pair of open-source streaming ASR and diarization models, the authors build four systems and evaluate multi-speaker accuracy, single-speaker accuracy degradation, memory footprint, and training complexity.

The result is a clarified design space with practical guidance on choosing the right architecture under different latency, compute, and memory constraints.

Original post →

More from Research

Research channel →