GPT-Live τ-Voice scores span 20 points across sources; backend model choice matters
rdesh26 · x · 2026-09-12
The author questions GPT-Live's several τ-Voice numbers: 86.2% on OpenAI's blog, 81.7% on the leaderboard, 67.9% on Artificial Analysis. Stochasticity explains only some of the 20-point gap; backend model choice matters a lot (see the Astra vs Sol difference). He recommends OpenAI's CRAWL/WALK/RUN eval harness for voice-agent evals. Part of the GPT-Live-1 dissection thread.
More from Models
- Grok Bot rolls out on Grok web app, xAI pitches it as your everything app — nima_owji · 2026-09-12
- John Schulman joins Dwarkesh podcast to debate RSI, RL, and AGI timelines — ZhongRuiqi · 2026-09-12
- Signal65: token prices rise but cost per finished AI task stays flat as agent accuracy jumps — ryanshrout · 2026-09-12
- DeepSeek v4.1 Flash runs out of the box on six NVIDIA GPUs via vLLM on day 0, AMD lags — woosuk_k · 2026-09-12
- ARC-AGI cost collapse: DeepSeek-V4-Flash hits 87% at $0.021/task vs o3's $4,500 — inductionheads · 2026-09-12
- Open uncensored AI has demand, but distribution has been stuck for two years — Rumbleblak · 2026-09-12