Cartesia launches Sonic 3.6: 44-language TTS model with sub-90ms latency
emmanuelvivier · x · 2026-08-20
Cartesia has released Sonic-3.6, the latest version of its real-time text-to-speech model. It ranks #1 on both Artificial Analysis speech leaderboards (1283 Elo on Provider Voice, 1123 on Controlled Voice). Built on state space models instead of Transformers, the model features sub-90ms time-to-first-audio latency and supports 44 languages. It is available via a beta API with no open-source weights.
More from Models
- Nvidia prioritizes Nemotron open-source models to rival top global models — pstAsiatech · 2026-08-20
- Same GRPO recipe on three from-scratch LLMs yields wildly different results with no clean scale relationship — john_enev · 2026-08-20
- User complains Fable model repeatedly lies about work done — JoshuaJBouw · 2026-08-20
- Scaling laws for diffusion MoE released; LLaDA v2 rivals Qwen3 with 35% fewer tokens — jiqizhixin · 2026-08-20
- Open Source Models Write Better Than OpenAI/Anthropic, User Claims — JoshPurtell · 2026-08-20
- Uncensored ERP Review: Why Qwen3-235B-a22b Remains the Top Choice — redditaccountno6 · 2026-08-20