ττ-Bench: Best coding agent passes just 23.9% of realistic agent-building tasks
rohanpaul_ai · x · 2026-09-21
A new arXiv paper introduces ττ-Bench, a benchmark that makes agent construction the task itself: a developer agent receives real business records, a client holding requirements, a production API, a codebase to inherit, and cost/model limits.
Key findings:
- 53 tasks across four domains; agents are scored by deploying their customer-service agent against held-out simulated users.
- The strongest setup, Claude Opus 5 under Claude Code, passes only 23.9% of evaluations vs. an 82.2% expert-authored ceiling.
- Failures mirror human dev pitfalls: shallow queries instead of deep understanding of records, almost no client communication, and too little experimentation with architecture and serving spend.
The authors conclude that better code generation alone won't fix agents' shallow requirement comprehension.
Related event: ττ-bench Exposes Coding Agents: Top Model Passes Only 23.9%(2 posts)→
More from coding & agent
- Google open-sources ARTEMIS, letting AI agents control a real phone end to end — dr_cintas · 2026-09-21
- Anyone's agents actually making money? A dev's reality check on x402 agent payments — thranduilsson · 2026-09-21
- Tsinghua's DiffuTester generates unit tests with diffusion LLMs 2-3x faster — jiqizhixin · 2026-09-21
- Turning a Linux desktop into an agent workspace: Claude, Codex and Hermes on Omarchy — Teknium · 2026-09-21
- Developer builds jev, an interpreter that reasons over plain-English facts and rules, inspired by Geoffrey Litt — narphorium · 2026-09-21
- evmscope MCP server ships 20 blockchain tools for AI agents across 5 EVM chains — modelcontextprotocol · 2026-09-21