Microsoft Releases VibeVoice-ASR-Streaming-7B for Streaming Chinese/English ASR

microsoft · hf · 2026-09-02

Microsoft released VibeVoice-ASR-Streaming-7B on Hugging Face, a 7B-parameter streaming automatic speech recognition model supporting speech-to-text and transcription in Chinese and English. It is built on the transformers architecture with weights provided in safetensors format.

Original post →

More from Multimodal

Multimodal channel →