Microsoft releases VibeVoice-ASR-Streaming-7B open streaming speech recognition model

Acceptable-Cycle4645 · reddit · 2026-09-03

Microsoft has released VibeVoice-ASR-Streaming-7B on Hugging Face, an open-weights streaming speech recognition model. As the ASR variant of the VibeVoice family, it targets real-time transcription use cases, and its 7B scale makes it notably large among open ASR models — worth a look for developers needing low-latency streaming transcription.

Related event: Microsoft Open-Sources VibeVoice-ASR-Streaming-7B Streaming Speech Recognition Model(2 posts)→

Original post →

More from Multimodal

Multimodal channel →