GPT-Live-1 completes 83.6% of voice-agent tasks on first try vs 45.7% for GPT-Realtime-2.1
OpenAIDevs · x · 2026-09-11
OpenAI's dev account shared benchmarks for GPT-Live-1, its production-focused voice model. Paired with GPT-6 Astra at medium reasoning effort, it completed 83.6% of airline, retail and telecom support tasks on Tau3 at first attempt, versus 45.7% for GPT-Realtime-2.1. On TauBanking (finding info in banking docs and using tools to resolve requests) the pairing scored 38.1%. On conversational dynamics, it hit 97.3% on Artificial Analysis's benchmark evaluating turn-taking, interruptions and backchannels, plus results on Full Duplex Bench v1.5.
Related event: OpenAI launches GPT-Live-1 voice model on API(22 posts)→
More from Models
- Microsoft Patches Record 974 Vulnerabilities, Mostly Found by AI — Distinct-Question-16 · 2026-09-11
- DeepSeek V4.1 Flash tops Vals open-weight index at $0.30 per test, with the smallest skills gap — teortaxesTex · 2026-09-11
- Do You Really Need Flagship Models? Dev Argues Medium Effort Covers 80% of Coding — iamaliveix · 2026-09-11
- OpenAI appears to be quietly rolling out managed Agents on its platform — testingcatalog · 2026-09-11
- 30B Open Model OpenResearcher Beats GPT-4.1 on BrowseComp-Plus — TheZachMueller · 2026-09-11
- Surge AI evals: Claude Fable 5.1 leads at 68.7, Gemini 3.8 Flash jumps 12 points on frontier math — echen · 2026-09-11