OpenAI ships GPT-Live-1 in the API: one model reasons across audio, no more STT-LLM-TTS chaining

pbbakkum · x · 2026-09-11

OpenAI shipped GPT-Live-1 in the API. Peter Bakkum, MTS on the realtime APIs, explains the shift: instead of chaining separate speech-to-text, LLM, and text-to-speech models, GPT-Live-1 reasons across incoming and outgoing audio as one continuous interaction. That matters because VAD and turn detection have plagued voice agents with awkward pauses, interruptions, dropped context, and false responses to background noise. The model is built around a fluid, turnless interaction model, and alpha testers report conversations that feel far closer to talking with a person. Several developers argue GPT-Live in the API is a bigger deal than Gemini Astra for voice agent integrations.

Related event: OpenAI Launches GPT-Live-1 Speech Model with Full-Duplex API(30 posts)→

Original post →

More from Models

Models channel →