MSL launches Muse Voice Transcribe, a streaming audio model with real-time ASR and diarization
bowenc0221 · x · 2026-09-02
MSL has launched Muse Voice Transcribe, its first streaming audio perception model, performing ASR, diarization, and endpointing all in real time. It supports hour-long audio, 20+ speakers, multilingual input with seamless code-switching, and contextual biasing. The team describes multimodality as the core interaction layer between humans and AI, positioning Muse Voice Transcribe as the first milestone in bringing its real-time voice interaction models to the public, with continued improvements and a full voice interaction system on the roadmap.
More from Multimodal
- vLLM-Omni Renders MiniMax H3 Faster Than Playback: 10.1s Video in 8.7s — vllm_project · 2026-09-02
- vLLM + FastVideo achieve faster-than-playback video generation using MiniMax H3 — vllm_project · 2026-09-02
- NVIDIA Explains Why DLSS 5 Needs Generation to Break Reconstruction's Ceiling — ctnzr · 2026-09-02
- SenseNova U1.5 Lite Open Source: Handles Complex Prompts & Native 4K — HeyToha · 2026-09-02
- "The Last Rain King" walking through a flooded city, made with Midjourney v8.2 — tisch_eins · 2026-09-02
- Help: Flux2-dev ControlNet Broken in ComfyUI 0.30.1 — External-Orchid8461 · 2026-09-02