Meta's Muse Voice Transcribe tops Pipecat benchmark with lowest semantic WER for voice agents
bowenc0221 · x · 2026-09-15
Meta released Muse Voice Transcribe, a streaming speech-to-text model built for real-time voice agents, combining accuracy and low latency with diarization, keyword biasing, endpointing, and mixed-language transcription.
Pipecat v1.9.0 adds support for it. The open-source Pipecat STT benchmark measures "semantic word error rate" across 1,000 speech fragments—a metric that ignores transcription differences irrelevant to LLM understanding and tracks real pipeline accuracy better than standard WER. Muse achieves the lowest semantic WER of any model tested so far.
More from coding & agent
- inclusionAI's LLaDA-UI: 16.7B MoE diffusion VLM for GUI agents — inclusionAI · 2026-09-15
- Andrew Chen: Agents gave me three messy inboxes to check — andrewchen · 2026-09-15
- The detective novel thought experiment: why hidden states beat tokens for model handoff — CShorten30 · 2026-09-15
- oh-my-pi open-sources agent context compaction design docs, now at 31k GitHub stars — nirmal_dist · 2026-09-15
- RAFT v3.1 open-sources retrieval-augmented fine-tuning to clone human personas — jessi_cata · 2026-09-15
- 18 sub-agents cut to 1: Polylane argues sub-agents are just wrong — zeeg · 2026-09-15