Microsoft releases VibeVoice-ASR-Streaming-7B open streaming speech recognition model
Acceptable-Cycle4645 · reddit · 2026-09-03
Microsoft has released VibeVoice-ASR-Streaming-7B on Hugging Face, an open-weights streaming speech recognition model. As the ASR variant of the VibeVoice family, it targets real-time transcription use cases, and its 7B scale makes it notably large among open ASR models — worth a look for developers needing low-latency streaming transcription.
More from Multimodal
- MiniMax H3 lacks micro detail in single images — ComfyUI workflow help — Ok-Brain-5729 · 2026-09-03
- SolarWM: open data and scalable training for long-horizon video world models — CUHKSZ · 2026-09-03
- Creator crafts 5-minute Odyssey film entirely in Grok Imagine for xAI's $100K contest — Kyrannio · 2026-09-03
- World's first fully AI video streaming service launches, built on MiniMax h3 max turbo — jfischoff · 2026-09-03
- Muse Spark 1.3 generates a Minecraft clone from a single prompt — and the result holds up — alexandr_wang · 2026-09-03
- Scale CEO Praises Muse Spark 1.3's Interactive 3D Japanese Garden Scene as Huge Leap — alexandr_wang · 2026-09-03