'What must be orchestrated today is model behavior tomorrow': voice pipelines get absorbed into models
morqon · x · 2026-08-25
Quoting "what must be orchestrated today is model behavior tomorrow," the post cites Prash Mittal's thread on voice agents: 2 years ago they were chained systems — VAD for turn ends, ASR, an LLM producing responses and tool calls, TTS — with developers owning orchestration like turn-taking, latency and interruptions.
Models like gpt-realtime have since swallowed most of that pipeline: a single model natively accepts and produces audio, with parallel streams for transcripts and tool calls. The outer conversation "protocol" loop remains — the clearest example of "pushing things left" into the model.
More from AGI Musings
- How to use AI agents to simulate your career future based on history — lxfater · 2026-08-25
- Microfiction: Power crash makes general intelligence momentarily free — voooooogel · 2026-08-25
- Why is prompting like a casino? Reflections on creativity, games, and AI — Dimillian · 2026-08-25
- The DNA of Next-Gen Agents: Separate Intelligence from Continuity — KL_AIC · 2026-08-25
- AI: Democratizing Knowledge or Becoming the Ultimate Gatekeeper? — Chemical-Escape-5034 · 2026-08-25
- Palantir CEO: General knowledge is now free, only specific skills hold value — r0ck3t23 · 2026-08-25