Cartesia founder: voice AI's hard part is continuous state, not speech
rohanpaul_ai · x · 2026-09-22
- The assumption that one model architecture works equally well for thinking and interacting is questionable. Voice models must track a conversation as the sequence keeps growing, not just produce correct answers.
- The author argues the right abstraction for voice AI may be continuous state rather than speech itself.
- Cartesia's founder explains they entered the space from sequence modeling (not speech research), which is why the company bet heavily on state space models. Full interview on "The Neon Show" YouTube channel.
Related event: Cartesia Founder: Voice AI's Challenge Is Continuous State, Not Speech(2 posts)→
More from Models
- Grok 4.7 comparison clip shows major gains in 3D modeling and game physics over 4.6 — belce_dogru · 2026-09-22
- Hands-on: TypeSafe's tiny Jev classifier goes head-to-head with Claude Haiku/Sonnet/Opus on text understanding — vesko_st · 2026-09-22
- Hand-built anti-memorization benchmark: tiny Jev classifier matches Sonnet at ~150x lower cost — vesko_st · 2026-09-22
- HLE leaderboard: Grok 4.7 at #21 while Gemini 3.8 and Muse 1.3 lead by a margin — himanshustwts · 2026-09-22
- Grok 4.7's Terminal-Bench 4.0 coding score is 'horrendous', falling far behind OpenAI and Anthropic — daniel_mac8 · 2026-09-22
- Grok 4.7 Hits Vercel AI Gateway with 500K Context and 40% Off Until Sept 27 — cramforce · 2026-09-22