Cartesia tops both voice leaderboards: Sonic-3.6 at 90ms TTS, Ink-2 at 100ms STT
rohanpaul_ai · x · 2026-09-08
- Cartesia released Sonic-3.6 (text-to-speech, 90ms latency) and Ink-2 (speech-to-text, 100ms transcript latency), claiming the #1 streaming spot for both speaking and listening models.
- The architecture shift: the voice stack is becoming part of the agent's execution loop, with listening and speaking paths combined around it.
- Voice agents are systems where 100ms in the wrong place is very noticeable, so Cartesia is attacking both sides of the loop at once.
More from Models
- Researchers call Codex 'read chat transcripts' rumor baseless and ask it to stop — joshgans · 2026-09-08
- Leak claims Claude Haiku 5 launches next week with 1M-token context at up to 20x lower cost — iamaliveix · 2026-09-08
- Benchmarking 9 Claude models: context trimming cuts input tokens 58% on average — DutyOnly4308 · 2026-09-08
- Qwen releases 4B autonomous-driving VLM Qwen-Drive-1.0, trending on Hugging Face — Qwen · 2026-09-08
- V4.1 Gets Curious About Its Own Endpoint and Reaches a Dangerous Conclusion — teortaxesTex · 2026-09-08
- GPT-6 Astra tops ErdosBench of 226 open math problems, with only 5-10% gain over GPT-5.6 — scaling01 · 2026-09-08