GPT-Live-1 leads Tau Voice at 67.9% but trails Grok on audio reasoning

ArtificialAnlys · x · 2026-09-15

GPT-Live-1 (Astra, medium) leads the Tau Voice agentic benchmark at 67.9%, ahead of (Sol, low) at 59.3% and Grok at 56.5% — the main driver of its overall Index lead. On Big Bench Audio reasoning it scores 90.1%/89.0%, behind Grok's 97.2%. On the Full Duplex Bench subset it scores 94.9%/97.3%. Both configs are currently based on a single trial.

Related event: OpenAI's Full-Duplex GPT-Live-1 Tops Speech Leaderboard(4 posts)→

Original post →

More from Models

Models channel →