Sierra's τ^τ-bench: Humans Hit 82.2% While Best Model Claude Opus 5 Manages Only 23.9%

zainhas · x · 2026-09-09

Sierra's new tool-calling benchmark τ^τ-bench (hyper-tau-bench) exposes a massive gap between humans and frontier models on long-horizon tool-use tasks:

Community submissions of harness + builder-model configurations are accepted via pull request.

Related event: Sierra Launches τ^τ-bench, Humans Far Outpace Top AI Models(2 posts)→

Original post →

More from Models

Models channel →