ττ-bench Exposes Coding Agents: Top Model Passes Only 23.9%

A new arXiv benchmark, ττ-bench, tests agents as realistic client projects with scattered docs, an API, existing code, and a budget; the best coding agent passed only 23.9% of the customer-service agent construction tasks.

2026-09-21 ~ 2026-09-21 · 2 related posts