Meta Releases Muse Voice Transcribe: Real-Time Streaming Speech Model with Diarization
bowenc0221 · x · 2026-09-02
Meta has launched Muse Voice Transcribe, its first streaming audio perception model combining ASR, speaker diarization, and endpointing. It supports 70+ languages (25 at launch), hour-long sessions, 20+ speakers, and seamless code-switching. The model uses reinforcement learning to achieve 'adaptive delay,' waiting longer only for difficult words, establishing a new pareto frontier for speed/accuracy trade-offs. It is available via Meta Model API, Meta AI for Mac, and Muse Code.
More from Models
- Anthropic releases Claude 5.1 for coding and knowledge work — kieranklaassen · 2026-09-02
- Every's Vibe Check: Anthropic's Fable 5.1 Reclaims the Coding Crown — every · 2026-09-02
- Hands-on: Claude 5.1 shows "monster" coding capabilities, rebuilt app from one prompt — every · 2026-09-02
- Dev review: Claude 5.1 delights users with human-like collaboration feel — every · 2026-09-02
- Fable 5.1 benchmarks show insane jumps on Terminal/Science — kimmonismus · 2026-09-02
- Anthropic Fable 5.1 is now available in Claude Code v2.1 — BLUECOW009 · 2026-09-02