Meta's Muse Voice Transcribe tops Pipecat benchmark with lowest semantic WER for voice agents

bowenc0221 · x · 2026-09-15

Meta released Muse Voice Transcribe, a streaming speech-to-text model built for real-time voice agents, combining accuracy and low latency with diarization, keyword biasing, endpointing, and mixed-language transcription.

Pipecat v1.9.0 adds support for it. The open-source Pipecat STT benchmark measures "semantic word error rate" across 1,000 speech fragments—a metric that ignores transcription differences irrelevant to LLM understanding and tracks real pipeline accuracy better than standard WER. Muse achieves the lowest semantic WER of any model tested so far.

Original post →

More from coding & agent

coding & agent channel →