Cartesia launches Sonic 3.6: 44-language TTS model with sub-90ms latency

emmanuelvivier · x · 2026-08-20

Cartesia has released Sonic-3.6, the latest version of its real-time text-to-speech model. It ranks #1 on both Artificial Analysis speech leaderboards (1283 Elo on Provider Voice, 1123 on Controlled Voice). Built on state space models instead of Transformers, the model features sub-90ms time-to-first-audio latency and supports 44 languages. It is available via a beta API with no open-source weights.

Original post →

More from Models

Models channel →