Sierra open-sources hyper-τ-bench, a benchmark testing if coding agents can build agents
karthik_r_n · x · 2026-09-10
Sierra released and open-sourced hyper-τ-bench (published as τ^τ-bench), a new long-horizon evaluation measuring whether models can construct agents, not just act as one.
- Setup: a developer agent works in a sandboxed workspace with a simulated business's records and an on-demand simulated client; it must recover the spec from scattered evidence (handbooks, support tickets, spreadsheets, frontline know-how), design the architecture, and turn business actions into tools until a working customer-service agent is delivered.
- Adversarial details: the client's REST API may be subtly defective, forcing the agent to decide whether a bug is in the spec or the code; the delivered agent must run on a fixed model menu within a per-conversation cost budget.
- Scoring: the finished agent is deployed against simulated production traffic using fully verifiable τ-bench-style tests the developer never saw while building.
Context: τ-bench, launched in 2024, measured whether models could act as reliable customer-service agents; hyper-τ-bench targets the next layer — who builds the agent, increasingly the models themselves.
More from coding & agent
- Dev says Codex's Build macOS Apps plugin one-shot fixes what other tools struggled with — Dimillian · 2026-09-10
- Astra Ultra PRs an optimization to karpathy's minGPT, cutting parameter visits 422 to 77 — Kuprel · 2026-09-10
- High-signal evals beat harness tweaks beat post-training: a layered agent methodology — abeirami · 2026-09-10
- Tradooor Mirror mirrors leading traders' consensus positioning into Hyperliquid perps — templecrash · 2026-09-10
- Cognition's Devin agents factor RSA-260, set new record; RSA-1024 cost pegged at ~$30M — StefanoGogioso · 2026-09-10
- Harmonic's Aristotle: an AI agent that proves software correct with machine-checked proofs — satnam6502 · 2026-09-10