TerminalWorld: Top AI Agents Pass Only 62.5% of Real Terminal Tasks
jiqizhixin · x · 2026-07-03
UCL, Nanjing University, and Tencent jointly introduced TerminalWorld, a benchmark of 1,530 tasks across 18 workflows built by reverse-engineering real terminal recordings. The current optimal agent achieved a pass rate of only 62.5%, and its scores showed extremely low correlation with existing benchmarks, exposing the capability bottlenecks of AI agents in real terminal environments. The paper is available on arXiv, with code and the project page open-sourced.
More from coding & agent
- Tweaked orchestration skill turns agents into self-policing workflow — pvncher · 2026-07-27
- A practical map of 11 protocols in the modern AI agent stack — TheTuringPost · 2026-07-27
- Qwen Code nightly adds Goal v3 orchestration and workspace channel controls — qwen-code-ci-bot · 2026-07-27
- NVIDIA says Nemotron 3 Ultra hit 97.1% on agentic RTL chip-design tasks — NVIDIAAI · 2026-07-27
- Tokyo Agent Forge hackathon shipped production-ready AI agents in one day — DavidBennett__ · 2026-07-27
- Long-running agents will need immutable event logs, this thread argues — sebpaquet · 2026-07-27