Meta releases Muse Voice: SOTA real-time speech-to-text model for 70+ languages
bowenc0221 · x · 2026-09-02
Meta released Muse Voice Transcribe, its first real-time audio perception model. It achieves SOTA performance in streaming speech-to-text, supports 70+ languages, and natively handles speaker diarization and endpointing.
Related event: Meta Launches Muse Voice Transcribe, Claiming SOTA Streaming ASR(10 posts)→
More from Multimodal
- World Labs Unveils Atlas: A Native Multimodal World Model — gowthami_s · 2026-09-02
- AI generated images and panos look amazing — gowthami_s · 2026-09-02
- How I made a 4-minute AI film collaborating with agents and free tools — LudovicCreator · 2026-09-02
- MiniMax H3 Max: 10x Faster Video Generation with Frontier Quality — isidentical · 2026-09-02
- New models: Fable 5.1, Mythos 5.1, Atlas — himanshustwts · 2026-09-02
- Creator demos Atlas model with camera conditioning feature — gowthami_s · 2026-09-02