Cartesia explains why benchmarking TTS is extremely hard: no single number captures voice quality

saranormous · x · 2026-09-16

Voice AI company Cartesia shares why rigorous model evals matter for TTS: quality is multidimensional — naturalness, prosody, pauses, speaker consistency — and no single benchmark number can capture it, so there is no shortcut to benchmarking speech. Sarah notes she and krandiash have been discussing voice evals for two years, pointing to Cartesia's investment in this practice as hard-won wisdom.

Original post →

More from Multimodal

Multimodal channel →