Microsoft's VibeVoice-ASR-Streaming: first LLM-based streaming speaker-attributed ASR, 1.5B/7B open-sourced
realmrfakename · x · 2026-09-04
A Microsoft team (Li Dong, Furu Wei, et al.) released the VibeVoice-ASR-Streaming technical report, one of the first LLM-based end-to-end approaches to streaming speaker-attributed ASR. It interleaves fixed-size audio chunks, small lookahead audio, and previous text, producing 'who said what' as speech arrives without a separate diarization stage. The 7B model achieves the lowest average WER/CER across five eval sets and best-or-tied speaker attribution in 12 of 13 settings. 1.5B and 7B weights plus inference code are released.
Related event: Microsoft Open-Sources VibeVoice Streaming ASR Model(6 posts)→
More from Multimodal
- User combines GPT-6 Astra with a modified fal H3 Max Director in new video prototype — OdinLovis · 2026-09-04
- Ideogram's Painful bbox/ Prompting Made the Minimax-H3 Transition Easy — Nimblecloud13 · 2026-09-04
- Higgsfield demos text-to-3D pipeline: GPT-6 Astra vibes-codes an Oval Office scene in Blender — jxnlco · 2026-09-04
- First Video Tutorial: Building Character Sheets in Krea 2 for MiniMax-H3 — solomars3 · 2026-09-04
- Stolen Texture: a GPT Image 2 prompt that spreads product material to everything it touches — aziz4ai · 2026-09-04
- YarnGPT: a Nigerian TTS filling the accent gap ElevenLabs leaves open — saheedniyi_02 · 2026-09-04