TerminalWorld: Top AI Agents Pass Only 62.5% of Real Terminal Tasks
jiqizhixin · x · 2026-07-03
UCL, Nanjing University, and Tencent jointly introduced TerminalWorld, a benchmark of 1,530 tasks across 18 workflows built by reverse-engineering real terminal recordings. The current optimal agent achieved a pass rate of only 62.5%, and its scores showed extremely low correlation with existing benchmarks, exposing the capability bottlenecks of AI agents in real terminal environments. The paper is available on arXiv, with code and the project page open-sourced.
More from coding & agent
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- Steal this idea: prompt-to-hardware where agents assemble custom devices — paraschopra · 2026-09-11
- Model Is the Least Interesting Part: A Guide to Six Core AI Architectures from RAG to Multi-Agent — goyalshaliniuk · 2026-09-11
- Non-coder builds layered memory architecture: 20k tokens tracks a year of agent conversations — matteoianni · 2026-09-11