Sierra's τ^τ-bench: Humans Hit 82.2% While Best Model Claude Opus 5 Manages Only 23.9%
zainhas · x · 2026-09-09
Sierra's new tool-calling benchmark τ^τ-bench (hyper-tau-bench) exposes a massive gap between humans and frontier models on long-horizon tool-use tasks:
- Humans score 82.2%, while the best model, Claude Opus 5 (Claude Code + max), reaches only 23.9%;
- #2 and #3 are GPT-5.6-sol (22.0%) and GPT-5.6-terra (18.0%), both on Codex xhigh;
- Kimi K3 takes #4 and #5: OpenCode max scores 18.0% vs Kimi Code max at 17.9% — the same model measured differently across harnesses and dates, raising questions about benchmark stability;
- Claude Sonnet 5 ranks 6th at 14.9%;
- The banking sub-benchmark is the hardest for both models and humans; per-scenario breakdowns (airline/retail/telecom/banking) are viewable on the leaderboard.
Community submissions of harness + builder-model configurations are accepted via pull request.
Related event: Sierra Launches τ^τ-bench, Humans Far Outpace Top AI Models(2 posts)→
More from Models
- DeepSeek routes V4 Pro requests to faster, cheaper V4.1 Flash, hints V4.1 Pro is coming — zephyr_z9 · 2026-09-09
- Fable's Test Persona Is Named Alex, and She's Nearly Ready for Users to Break — RileyRalmuto · 2026-09-09
- DeepSeek reportedly admits V4-Pro pretraining misstep; V4.1-Flash due Sept 10 — teortaxesTex · 2026-09-09
- AI 2027 was mocked as too fast — GPT-6 Astra just landed on its curve — Confident_Salt_8108 · 2026-09-09
- Anthropic admits internal models used in math proof, sparking transparency row — teortaxesTex · 2026-09-09
- shuding Responds to Highlighting Benchmark Backlash: Prism.js Label Mismatch Explains Gap — shuding · 2026-09-09