Meta Releases Muse Voice Transcribe: Real-Time Streaming Speech Model with Diarization

bowenc0221 · x · 2026-09-02

Meta has launched Muse Voice Transcribe, its first streaming audio perception model combining ASR, speaker diarization, and endpointing. It supports 70+ languages (25 at launch), hour-long sessions, 20+ speakers, and seamless code-switching. The model uses reinforcement learning to achieve 'adaptive delay,' waiting longer only for difficult words, establishing a new pareto frontier for speed/accuracy trade-offs. It is available via Meta Model API, Meta AI for Mac, and Muse Code.

Related event: Meta Superintelligence Labs Launches Muse Voice Transcribe, Its First Real-Time Audio Perception Model(8 posts)→

Original post →

More from Models

Models channel →