FULL STORY

GPT-Live-1 Debut and the Benchmark Caveat

OpenAI's full-duplex speech model GPT-Live-1 topped voice benchmarks on debut, but analysts noted its scores relied on GPT-6 Astra as backend reasoning model.

2026-09-15 ~ 2026-09-15 · 2 episodes · 8 posts

Episode 1 · OpenAI's Full-Duplex Speech Model GPT-Live-1 Debuts atop Voice Leaderboard (2026-09-15, 6 posts)

OpenAI has released GPT-Live-1, a full-duplex speech-to-speech model. According to Artificial Analysis, it debuted at the top of the Speech to Speech Index with a score of 81.5, edging out Grok Voice Think Fast 2.0 at 81.3, while its per-hour audio cost is significantly lower than the previous-generation Realtime model. However, category-level evaluations show it does not lead across the board on audio reasoning and user preference, with Grok still ahead on some dimensions. The model can also delegate reasoning to a backend text model.

Confirmed

  • On the Artificial Analysis Speech to Speech Index, GPT-Live-1 ranks first with 81.5, versus 81.3 for Grok Voice Think Fast 2.0
  • Cost estimate: on a fixed 40-question Big Bench Audio subset, including voice session fees and backend text model token costs, the (Astra, medium) configuration costs about $5 per hour of input audio, with an overall range of $4.47-5.83, far below the previous-generation Realtime
  • On the agentic benchmark Tau Voice, (Astra, medium) leads with 67.9%, followed by (Sol, low) at 59.3% and Grok at 56.5%, the main driver of its top overall ranking
  • It still trails Grok on Big Bench Audio speech/audio reasoning
  • Speech Agent Arena preference leaderboard: GPT-Live-1 (Sol, low) ranks 3rd with 1,053 Elo and a 90.9% task success rate, while (Astra, medium) ranks 4th with 1,048 Elo; Grok leads on task success rate

Why it matters

  • GPT-Live-1 demonstrates an architecture where speech models delegate reasoning to backend text models, yielding advantages in cost and agentic task capability
  • Topping the leaderboard while trading wins across subcategories suggests the speech model landscape remains open, with Grok retaining strengths in audio reasoning and user preference

Episode 2 · GPT-Live-1 Benchmark Scores Depend on Backend Model (2026-09-15, 2 posts)

The founder of open-source voice agent platform Dograh notes that GPT-Live-1's benchmark results were achieved with GPT-6 Astra as the backend, so real-world performance depends on developers' own inference models.