GPT-Live τ-Voice scores span 20 points across sources; backend model choice matters

rdesh26 · x · 2026-09-12

The author questions GPT-Live's several τ-Voice numbers: 86.2% on OpenAI's blog, 81.7% on the leaderboard, 67.9% on Artificial Analysis. Stochasticity explains only some of the 20-point gap; backend model choice matters a lot (see the Astra vs Sol difference). He recommends OpenAI's CRAWL/WALK/RUN eval harness for voice-agent evals. Part of the GPT-Live-1 dissection thread.

Original post →

More from Models

Models channel →