GPT-Live-1 leads Tau Voice at 67.9% but trails Grok on audio reasoning
ArtificialAnlys · x · 2026-09-15
GPT-Live-1 (Astra, medium) leads the Tau Voice agentic benchmark at 67.9%, ahead of (Sol, low) at 59.3% and Grok at 56.5% — the main driver of its overall Index lead. On Big Bench Audio reasoning it scores 90.1%/89.0%, behind Grok's 97.2%. On the Full Duplex Bench subset it scores 94.9%/97.3%. Both configs are currently based on a single trial.
Related event: OpenAI's Full-Duplex GPT-Live-1 Tops Speech Leaderboard(4 posts)→
More from Models
- OpenAI researcher explains why lab staff are suddenly scared about AI progress — socoolandawesome · 2026-09-15
- Local Qwen loops and forgets in coding agents while Claude Code just works — tlpta · 2026-09-15
- Rumor: frontier lab training new internal model since Aug 28 with 'unprecedented' math performance — rapha_gl · 2026-09-15
- User Slams 'Puritanical' AI Censorship Over Refused Cuddling Image — MeroveeFrancSalien · 2026-09-15
- Prefill and Decode: why asking an LLM for three takeaways from a long document still takes minutes — dotey · 2026-09-15
- Leak claims xAI trails OpenAI and Anthropic by roughly 6-12 months — sachinmaya1980 · 2026-09-15