Meta Releases Muse Voice Transcribe: Real-Time ASR with 20+ Speaker Diarization

AIatMeta · x · 2026-09-02

Meta has released Muse Voice Transcribe, the first real-time audio perception model from Meta Superintelligence Labs. It features streaming ASR with ultra-low latency, diarization supporting over 20 speakers simultaneously, and native multilingual support for 25+ languages with seamless code-switching. The model also leverages keyword and context biasing for improved accuracy and is available via Meta Model API, Meta AI for Mac, and Muse Code.

Related event: Meta Launches Muse Voice Transcribe, Claiming SOTA Streaming ASR(10 posts)→

Original post →

More from Models

Models channel →