Microsoft Releases VibeVoice-ASR-Streaming-7B for Streaming Chinese/English ASR
microsoft · hf · 2026-09-02
Microsoft released VibeVoice-ASR-Streaming-7B on Hugging Face, a 7B-parameter streaming automatic speech recognition model supporting speech-to-text and transcription in Chinese and English. It is built on the transformers architecture with weights provided in safetensors format.
More from Multimodal
- Video generation has matured: you can now direct models instead of prompting and hoping — MilitantAI · 2026-09-03
- Video gen has moved from prompting to directing — and product UX isn't ready — Kyrannio · 2026-09-03
- World Labs unveils Atlas, a new video generation model — mildlyphd · 2026-09-03
- Runway's Solaris: an interface world model that renders apps pixel by pixel — umpherj · 2026-09-03
- VideoDeltaNet open-sources hybrid attention that speeds up MiniMax H3 video generation up to 90x — realmrfakename · 2026-09-03
- jjk-explain turns any concept into a Jujutsu Kaisen-style explainer video with one Claude Code command — teortaxesTex · 2026-09-03