OpenAI ships GPT-Live-1 in the API: one model reasons across audio, no more STT-LLM-TTS chaining
pbbakkum · x · 2026-09-11
OpenAI shipped GPT-Live-1 in the API. Peter Bakkum, MTS on the realtime APIs, explains the shift: instead of chaining separate speech-to-text, LLM, and text-to-speech models, GPT-Live-1 reasons across incoming and outgoing audio as one continuous interaction. That matters because VAD and turn detection have plagued voice agents with awkward pauses, interruptions, dropped context, and false responses to background noise. The model is built around a fluid, turnless interaction model, and alpha testers report conversations that feel far closer to talking with a person. Several developers argue GPT-Live in the API is a bigger deal than Gemini Astra for voice agent integrations.
Related event: OpenAI Launches GPT-Live-1 Speech Model with Full-Duplex API(30 posts)→
More from Models
- Hands-on with GPT-6 Astra: stunning at 3D games, but two flaws keep it from daily driver — petergyang · 2026-09-11
- Only Muse Spark 1.3 and Fable 5.1 sit on the coding Pareto frontier — jyangballin · 2026-09-11
- Muse Spark 1.3 shines on brand-new, unoptimizable CursorBench 4.0 — jyangballin · 2026-09-11
- User slams Anthropic for blocking benign queries on ancient texts and recursive AI — NickPassig · 2026-09-11
- Codex cybersecurity work needs the Daybreak model to avoid safety-guardrail blocks — HankYeomans · 2026-09-11
- DeepSeek's new open-source model reportedly crushes GLM and Kimi at 4-10x lower prices — anselm · 2026-09-11