MSL launches Muse Voice Transcribe, a streaming audio model with real-time ASR and diarization

bowenc0221 · x · 2026-09-02

MSL has launched Muse Voice Transcribe, its first streaming audio perception model, performing ASR, diarization, and endpointing all in real time. It supports hour-long audio, 20+ speakers, multilingual input with seamless code-switching, and contextual biasing. The team describes multimodality as the core interaction layer between humans and AI, positioning Muse Voice Transcribe as the first milestone in bringing its real-time voice interaction models to the public, with continued improvements and a full voice interaction system on the roadmap.

Related event: Meta Superintelligence Labs Launches Muse Voice Transcribe, Its First Real-Time Audio Perception Model(9 posts)→

Original post →

More from Multimodal

Multimodal channel →