113 TTS eval papers mapped: WER rewards kill prosody, UTMOS doesn't travel, and MOS scores aren't comparable
Slight_Republic_4242 · reddit · 2026-09-09
A voice-agent team mapped 113 TTS evaluation papers onto a 50-leaf taxonomy while picking CI metrics, and found four things that changed their plans:
- WER as an optimization target eats prosody. A GRPO run trained on transcription-oriented signals (CER/NLL, arXiv 2509.18531) lowered error rates while prosody collapsed into monotone; adding speaker-similarity to the reward destabilized training further. WER+SECS is fine as a guardrail, bad as an objective.
- UTMOS doesn't travel. TTScore (2509.20485) showed UTMOS performs well on its training domain (VoiceMOS) but degrades on SOMOS enough that a metric never trained on MOS labels beats it — MOS-labelled data is too small to generalize.
- MOS numbers across papers are not comparable at all (Dagstuhl good-practices doc, 2503.03250); a true MUSHRA for TTS can't exist for lack of anchors, so vendor cross-quoting of MOS is noise.
- Nothing measures what production cares about: turn-taking, barge-in, endpointing latency, streaming-lookahead quality loss, and 8 kHz telephony degradation sit outside every TTS benchmark (relevant work like Full-Duplex-Bench, SPEARBench, FastTurn exists but isn't integrated). Everyone benchmarks clean 24 kHz read speech.
Six taxonomy leaves came back nearly empty: abbreviations, URLs/addresses/equations, syntactic complexity, code-switching, unseen linguistic structures, and accent — the first three carried almost solely by EmergentTTS-Eval.
More from Multimodal
- Full video prompt shared for making a cinematic 10-second travel montage with Gemini Omni — michaelrabone · 2026-09-09
- GPT Image 2.5 + MiniMax H3 combo turns a single reference image into 16-panel video — Hailuo_AI · 2026-09-09
- Cinematic travel montages made with Omni in Google Gemini, prompts shared — michaelrabone · 2026-09-09
- Prompt share: 'Retro Cel Fantasy' template for 80s hand-drawn animation-style images — azed_ai · 2026-09-09
- Full-AI pipeline turns Blender city into stylized video via GPT, DepthAnything and Hailuo — Hailuo_AI · 2026-09-09
- ChatGPT Images 2.5 generates website mockups with no broken text — aitrendz_xyz · 2026-09-09