Pipecat v1.9 adds Meta's Muse Voice Transcribe, the lowest semantic-WER STT model tested

solyarisoftware · x · 2026-09-12

The voice agent framework Pipecat shipped v1.9.0 with support for Meta's new streaming speech-to-text model, Muse Voice Transcribe.

The team maintains Pipecat STT benchmark, an open source test suite that measures STT performance inside a voice agent pipeline. It computes a "semantic word error rate" over 1,000 speech fragments — ignoring transcription differences that don't affect an LLM's understanding of user speech — which they find a better proxy for real accuracy than standard WER algorithms. Muse scored the best (lowest) semantic WER of any model they've tested.

They also measure latency as "time to final segment" of the transcription, a critical metric for voice agents.

Original post →

More from coding & agent

coding & agent channel →