Meta Launches Muse Voice Transcribe, Claiming SOTA Streaming ASR
Meta Superintelligence Labs officially launched Muse Voice Transcribe on September 2, its first real-time audio perception model, achieving SOTA on streaming speech-to-text (ASR). It marks Meta's entry from transcription into real-time audio perception, and is a flagship joint result after Scale AI joined Meta (Scale AI founder Alexandr Wang posted about it the same day).
Confirmed
- Core capabilities confirmed by Meta's official account @AIatMeta: low-latency real-time streaming ASR, speaker diarization (supporting 20+ speakers), and endpointing
- Trained on 70+ languages, with 25 supported at launch, including seamless multilingual code-switching, even within a sentence
- Handles long-form audio of over 1 hour
- Supports biasing by language, keywords, and context to improve recognition accuracy
- Available via Meta's model API with a zero data retention tier, and usable in the Meta desktop app and Muse Code
Why it matters
- Real-time ASR, speaker diarization, and endpointing are natively integrated in a single model, eliminating traditional pipeline stitching—directly useful for real-time translation, meeting transcription, and voice assistants
- 20+ speaker diarization and intra-sentence code-switching are common weak spots in similar models; Meta is clearly targeting real-world multilingual, multi-speaker conversations
- The zero data retention tier lowers privacy compliance barriers for enterprises, signaling Meta's push into commercial API services
2026-09-02 ~ 2026-09-02 · 10 related posts
Primary sources
- Meta Releases Muse Voice Transcribe: Real-Time ASR with 20+ Speaker Diarization — AIatMeta ·
- Scale AI Releases Muse Voice: SOTA Streaming Speech-to-Text Model — alexandr_wang ·
- Meta releases Muse Voice Transcribe with top streaming accuracy — ArtificialAnlys ·
- Meta releases Muse Voice: SOTA real-time speech-to-text model for 70+ languages — bowenc0221 · 2026-09-02
- Meta's Muse Voice Transcribe hits API with zero-data-retention tier — testingcatalog · 2026-09-02
- Meta's Muse Voice Transcribe Balances Speed and Accuracy with Adaptive Delay — AIatMeta · 2026-09-02
- [source] Meta Releases Muse Voice Transcribe: Real-Time ASR with 20+ Speaker Diarization — AIatMeta · 2026-09-02
- Meta Releases Muse Voice Transcribe: Real-Time Streaming Speech Model with Diarization — bowenc0221 · 2026-09-02
- MSL launches Muse Voice Transcribe, a streaming audio model with real-time ASR and diarization — bowenc0221 · 2026-09-02
- [source] Scale AI Releases Muse Voice: SOTA Streaming Speech-to-Text Model — alexandr_wang · 2026-09-02
- [source] Meta releases Muse Voice Transcribe with top streaming accuracy — ArtificialAnlys · 2026-09-02
2 near-duplicate retellings: bowenc0221 · rohanpaul_ai