OpenAI: GPT-Live-1 completes 83.6% of voice-agent tasks first try vs 45.7% for GPT-Realtime-2.1
OpenAIDevs · x · 2026-09-11
OpenAI's dev account published production benchmarks for GPT-Live-1: paired with GPT-6 Astra at medium reasoning effort, it completed 83.6% of Tau3 voice-agent tasks (airline/retail/telecom) on first attempt vs 45.7% for GPT-Realtime-2.1; scored 38.1% on TauBanking (document retrieval + tool use); and hit 97.3% on Artificial Analysis's Conversational Dynamics benchmark for turn-taking and interruption handling.
Related event: OpenAI launches full-duplex speech model GPT-Live-1 on API(26 posts)→
More from Models
- Anthropic says it halted plots to use its AI models to help develop bioweapons — austinc3301 · 2026-09-11
- Humans&AI launches Persimmon, first large-scale model to simulate how people talk and interact — gharik · 2026-09-11
- MTP dropped in new model's tech report: fp8 lookup tables and prime table sizes noted — stochasticchasm · 2026-09-11
- DeepSeek unveils V4.1-Flash: new encoder-decoder architecture with native vision — gaganghotra_ · 2026-09-11
- User suspects Ox Alpha and GLM 5.3 Flash outputs were secretly routed to Claude — liminal_bardo · 2026-09-11
- Dev switches daily driver to GPT-6 Astra Low: near-flagship quality, faster and cheaper — intellectronica · 2026-09-11