Microsoft's VibeVoice-ASR-Streaming: first LLM-based streaming speaker-attributed ASR, 1.5B/7B open-sourced

realmrfakename · x · 2026-09-04

A Microsoft team (Li Dong, Furu Wei, et al.) released the VibeVoice-ASR-Streaming technical report, one of the first LLM-based end-to-end approaches to streaming speaker-attributed ASR. It interleaves fixed-size audio chunks, small lookahead audio, and previous text, producing 'who said what' as speech arrives without a separate diarization stage. The 7B model achieves the lowest average WER/CER across five eval sets and best-or-tied speaker attribution in 12 of 13 settings. 1.5B and 7B weights plus inference code are released.

Related event: Microsoft Open-Sources VibeVoice Streaming ASR Model(6 posts)→

Original post →

More from Multimodal

Multimodal channel →