Sierra's New Tool-Calling Bench τ^τ-bench: Kimi K3 Scores Differ Two Days Apart, Banking Sub-bench Hardest
zainhas · x · 2026-09-09
Sierra released τ^τ-bench (hyper-tau-bench), a new benchmark for tool calling in multi-turn business scenarios.
The poster highlights two observations:
- Kimi K3 measured differently two days apart, and the coding harness (OpenCode vs Kimi Code) also changes results, suggesting the bench is sensitive to runtime setup;
- The banking sub-benchmark is the hardest, challenging for both models and humans.
Detailed leaderboard data is covered in the companion post.
Related event: Sierra Launches τ^τ-bench, Humans Far Outpace Top AI Models(2 posts)→
More from Models
- OpenAI may pause new Pro subscriptions as demand for Astra hits unprecedented levels — op7418 · 2026-09-09
- Veteran claims a trained intuition for sniffing out AI-written text — IndraVahan · 2026-09-09
- Tencent Hunyuan details Gander, an end-to-end full-duplex omni interaction agent — Tencent-Hunyuan · 2026-09-09
- Moonshot open-sources MoBA, a block attention mechanism it says is 16x faster for long context — gekobraa · 2026-09-09
- From Failing 2+2 to PhD-Level Math: A Timeline of AI's Three-Year Leap — haider1 · 2026-09-09
- Muse Spark 1.3 reportedly crushes tax agent bench at $0.32/task, 3x cheaper — zainhas · 2026-09-09