Cartesia's new voice model family draws praise; WER alone can't capture context-correct speech
buckymoore · x · 2026-09-16
Voice AI startup Cartesia explains how it benchmarks speech models: generate audio, transcribe with an ASR model, and compare to the original text via Word Error Rate (WER). But WER only tells whether the model said the right words — a sentence can be correct word-by-word yet wrong in context. Buckymoore praised the quality of Cartesia's latest model family and asked whether any other lab is scaling an alternative architecture with such results.
More from Models
- With retries and pooled selection, Qwen3.8 27B hits 92.04% on DeepSWE 1.1, ~18 pts above GPT-6 Astra — S_Conradi · 2026-09-16
- "Just output probability distributions, never hallucinate": AI safety claim gets mocked — inductionheads · 2026-09-16
- Chinese open models hit 53% of OpenRouter tokens, but closed models still dominate real adoption — ohlennart · 2026-09-16
- Niche AI Use Cases Keep Getting Absorbed Into General Models — samiramanabi · 2026-09-16
- Google reportedly building math-focused DeepThink variant, raw thoughts leak — PMinervini · 2026-09-16
- Player claims to find OpenAI GPT-6 "Astra" easter egg in Fallout 3 — imjustnewatai · 2026-09-16