Meta Releases MuseVoiceTranscribe: Real-time Audio Perception Model for 20+ Speakers
智东西 · wechat · 2026-09-02
Meta has released MuseVoiceTranscribe, its first real-time audio perception model, aiming to upgrade speech transcription into a real-time sensing layer for AI systems. Key capabilities include:
- Multilingual & Long Audio: Supports 70+ languages (25 verified), natively handles audio over 1 hour, and enables seamless code-switching (e.g., Chinese-English mixing).
- Multi-speaker Diarization: Natively supports diarization for 20+ speakers without post-processing, accurately identifying speaker turns.
- Streaming & Adaptive Latency: Uses an autoregressive multimodal architecture processing audio in 80ms chunks. It leverages Reinforcement Learning (RL) for "adaptive latency," allowing the model to dynamically decide listening duration per word to balance speed and accuracy optimally.
- Voice Activity Detection: Introduces special tokens to mark speech onset and endpoints.
The model is integrated into Meta AI, MuseCode, and MetaModelAPI, priced at approx. $3 per 1,000 minutes.
Related event: Meta Launches Muse Voice Transcribe, a Real-Time Audio Perception Model(14 posts)→
More from Multimodal
- Infinite AI livestream powered by MiniMax H3 generates faster than playback — FinanceYF5 · 2026-09-02
- AI can create 3D worlds with keyframed camera angles, boosting filmmaking — drfeifei · 2026-09-02
- TenStrip 10Eros Models: High-Gen Checkpoints with Built-in Turbo LoRA — Tight_Organization54 · 2026-09-02
- Job seeker showcases AI video reel built with Minimax H3 — PurzBeats · 2026-09-02
- LTX2.3 Generated Video: The Hatter Names All the Hats — Tokyo_Jab · 2026-09-02
- User showcases impressive video generation demo from Claude Fable 5.1 — zainhas · 2026-09-02