Cartesia's new voice model family draws praise; WER alone can't capture context-correct speech

buckymoore · x · 2026-09-16

Voice AI startup Cartesia explains how it benchmarks speech models: generate audio, transcribe with an ASR model, and compare to the original text via Word Error Rate (WER). But WER only tells whether the model said the right words — a sentence can be correct word-by-word yet wrong in context. Buckymoore praised the quality of Cartesia's latest model family and asked whether any other lab is scaling an alternative architecture with such results.

Original post →

More from Models

Models channel →