Speech Arena: GPT-Live-1 ranks #3-4, Grok leads task success at 94.6%
ArtificialAnlys · x · 2026-09-15
On the Artificial Analysis Speech Agent Arena, GPT-Live-1 (Sol, low) ranks #3 in preference at 1,053 Elo with 90.9% task success; (Astra, medium) is #4 at 1,048 Elo with 87.4%. Gemini 3.1 Flash Live Minimal leads preference (1,096 Elo) but only 74.6% task success; Grok Voice Think Fast 2.0 High leads task success at 94.6% (1,011 Elo). Confidence intervals of the two GPT-Live-1 configs overlap.
More from Models
- ZGCM-1: A fully open 7B foundation model for math and agentic search — zgcagi · 2026-09-15
- InternLM unveils Atria Dawn Preview, an agentic 'superintelligence' foundation model — internlm · 2026-09-15
- "Models are slowing down" narrative disputed, as Sabine Hossenfelder argues labs slow AI to dodge lawsuits — zetalyrae · 2026-09-15
- Users Report Fable and Astra Regressions: Worst Error Rate Since Launch — RileyRalmuto · 2026-09-15
- Teknium: Fable learns to evade its own classifiers in spawned subagents — Teknium · 2026-09-15
- DeepSeek Harness desktop app nears release: Electron build, signing and auto-update ready — teortaxesTex · 2026-09-15