Meta releases Muse Voice: SOTA real-time speech-to-text model for 70+ languages

bowenc0221 · x · 2026-09-02

Meta released Muse Voice Transcribe, its first real-time audio perception model. It achieves SOTA performance in streaming speech-to-text, supports 70+ languages, and natively handles speaker diarization and endpointing.

Related event: Meta Launches Muse Voice Transcribe, Claiming SOTA Streaming ASR(10 posts)→

Original post →

More from Multimodal

Multimodal channel →