'What must be orchestrated today is model behavior tomorrow': voice pipelines get absorbed into models

morqon · x · 2026-08-25

Quoting "what must be orchestrated today is model behavior tomorrow," the post cites Prash Mittal's thread on voice agents: 2 years ago they were chained systems — VAD for turn ends, ASR, an LLM producing responses and tool calls, TTS — with developers owning orchestration like turn-taking, latency and interruptions.

Models like gpt-realtime have since swallowed most of that pipeline: a single model natively accepts and produces audio, with parallel streams for transcripts and tool calls. The outer conversation "protocol" loop remains — the clearest example of "pushing things left" into the model.

Original post →

More from AGI Musings

AGI Musings channel →