GPT-Live-1 completes 83.6% of voice-agent tasks on first try vs 45.7% for GPT-Realtime-2.1

OpenAIDevs · x · 2026-09-11

OpenAI's dev account shared benchmarks for GPT-Live-1, its production-focused voice model. Paired with GPT-6 Astra at medium reasoning effort, it completed 83.6% of airline, retail and telecom support tasks on Tau3 at first attempt, versus 45.7% for GPT-Realtime-2.1. On TauBanking (finding info in banking docs and using tools to resolve requests) the pairing scored 38.1%. On conversational dynamics, it hit 97.3% on Artificial Analysis's benchmark evaluating turn-taking, interruptions and backchannels, plus results on Full Duplex Bench v1.5.

Related event: OpenAI launches GPT-Live-1 voice model on API(22 posts)→

Original post →

More from Models

Models channel →