Voice AI's hard part may be continuous state, not speech: Cartesia take
rohanpaul_ai · x · 2026-09-22
- The author challenges the assumption that one model architecture should serve both thinking and interacting: a voice model must keep up with a human while the sequence keeps growing, not just produce correct answers.
- He argues "speech" may not be the right abstraction for the hard part — "continuous state" could be — and points to Cartesia's approach; full discussion in his video on The Neon Show YouTube channel.
Related event: Cartesia Founder: Voice AI's Challenge Is Continuous State, Not Speech(2 posts)→
More from Models
- Grok 4.7 comparison clip shows major gains in 3D modeling and game physics over 4.6 — belce_dogru · 2026-09-22
- Hands-on: TypeSafe's tiny Jev classifier goes head-to-head with Claude Haiku/Sonnet/Opus on text understanding — vesko_st · 2026-09-22
- Hand-built anti-memorization benchmark: tiny Jev classifier matches Sonnet at ~150x lower cost — vesko_st · 2026-09-22
- HLE leaderboard: Grok 4.7 at #21 while Gemini 3.8 and Muse 1.3 lead by a margin — himanshustwts · 2026-09-22
- Grok 4.7's Terminal-Bench 4.0 coding score is 'horrendous', falling far behind OpenAI and Anthropic — daniel_mac8 · 2026-09-22
- Grok 4.7 Hits Vercel AI Gateway with 500K Context and 40% Off Until Sept 27 — cramforce · 2026-09-22