GPT-5.2 + Codex CLI Hits 63% on Terminal-Bench 2.0's 89 Hard CLI Tasks

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini, Boxuan Li, Harsh Raj, Ivan Bercovich, Lin Shi, Jeong Yeon Shin, Thomas Walshe, E. Kelly Buchanan, Junhong Shen, Guanghao Ye, Haowei Lin, Jason Poulos, Maoyu Wang, Marianna Nezhurina, Jenia Jitsev, Di Lu, Orfeas Menis Mastromichalakis, Zhiwei Xu, Zizhao Chen, Yue Liu, Robert Zhang, Leon Liangyu Chen, Anurag Kashyap, Jan-Lucas Uslu, Jeffrey Li, Jianbo Wu, Minghao Yan, Song Bian, Vedang Sharma, Ke Sun, Steven Dillmann, Akshay Anand, Andrew Lanpouthakoun, Bardia Koopah, Changran Hu, Etash Guha, Gabriel H. S. Dreiman, Jiacheng Zhu, Karl Krauth, Li Zhong, Niklas Muennighoff, Robert Amanfu, Shangyin Tan, Shreyas Pimpalgaonkar, Tushar Aggarwal, Xiangning Lin, Xin Lan, Xuandong Zhao, Yiqing Liang, Yuanli Wang, Zilong Wang, Changzhi Zhou, David Heineman, Hange Liu, Harsh Trivedi, John Yang, Junhong Lin, Manish Shetty, Michael Yang, Nabil Omi, Negin Raoof, Shanda Li, Terry Yue Zhuo, Wuwei Lin, Yiwei Dai, Yuxin Wang, Wenhao Chai, Shang Zhou, Dariush Wahdany, Ziyu She, Jiaming Hu, Zhikang Dong, Yuxuan Zhu, Sasha Cui, Ahson Saiyed, Arinbjörn Kolbeinsson, Jesse Hu, Christopher Michael Rytting, Ryan Marten, Yixin Wang, Alex Dimakis, Andy Konwinski, Ludwig Schmidt

cs.SE, cs.AI

2026-01-17

Stanford et al. release Terminal-Bench 2.0: 89 audited CLI tasks. GPT-5.2 with Codex CLI leads at 63%; best open-weight model hits 36%. Some tasks remain unsolved.

What problem this solves

Agent benchmarks sit at two unhelpful extremes. Synthetic environments miss paid work. Easier real-task suites have lost discriminating power at the frontier. The terminal sits in a useful middle: it is text, so language models can drive it, and it is where software engineering, scientific computing, security, and ML work actually get done. Cursor, Codex CLI, Claude Code, and Gemini CLI all operate by issuing shell commands. The paper cites Anthropic's claim that Claude Code had reached a $1B run-rate.

The missing piece was a hard exam that still looks like real work and can be graded automatically. Terminal-Bench puts agents in Dockerized shells and asks for billed-quality work: configure legacy systems, reimplement papers, fix compilers, rewrite COBOL in Python.

Method

Each task is a bundle: a natural-language instruction, a Docker image, hidden tests, a human-written oracle solution, and a time limit. Tests inspect the final container state only. They ignore the command trace. Any path is legal if the end state is right. Different scaffolds (the tool loop and prompting wrapped around a model) can therefore compete on the same items.

Tasks follow the Harbor format. Harbor is the accompanying eval runner. It can launch Claude Code, Codex CLI, OpenHands, Mini-SWE-Agent, and Terminus 2, a scaffold built for this paper. Terminus 2 exposes one tool, a headless terminal, and speaks only Bash. Commercial scaffolds are often tuned to a sibling model, so model quality and agent engineering get tangled. Terminus 2 is there to untangle them.

93 contributors submitted 229 items; 89 survived review by three experienced auditors, forming Terminal-Bench 2.0. Mean combined review time is about three hours per task. Gates include: the oracle must pass, a no-op dummy agent must fail, an LLM linter hunts spec bugs, an adversarial exploit agent looks for shortcuts, then strong models are rerun after merge for trajectory audit.

Software engineering is the largest slice at 26 tasks and still not a majority. On the subset with time estimates, experts finish 48.6% in under an hour and 47.3% in one hour to one day; juniors land 71.6% in the hour-to-day band, and 4.1% take more than a week. The heaviest item, fix-ocaml-gc, asks the agent to repair a failed OCaml GC optimization: about 24 hours for an expert, about 10 days for a junior.

The eval ran six agents across 16 frontier models. Each supported pair was run at least five times, 32,155 trials in total. Models with a reasoning-effort knob used the vendor default (medium for OpenAI and Anthropic). Closed models went through first-party APIs; open-weight models through Together.AI. Containers ran on Daytona, 32 to 100 in parallel.

Results

Figure 1 reports each model with the scaffold that maximized its score.

PairResolution rate
GPT-5.2 + Codex CLI63%
Claude Opus 4.5 + Terminus 258%
Gemini 3 Pro + Terminus 257%
Kimi K2 Thinking + Terminus 2 (best open-weight)36%
Smaller models15%

The top 13 slots are proprietary. Switching Codex CLI from GPT-5-Nano to GPT-5.2 raises resolution by 52%. Switching Gemini 2.5 Pro from OpenHands to Terminus 2 raises it by 17%. Picking the model usually moves the score more than picking the scaffold. A cluster of tasks is unsolved by every pair. System configuration, kernel-driver builds, and database migrations show up as common total failures.

A full run costs about one to a hundred dollars, depending on price. Most trials finish in under 20 minutes. Outliers reach two hours, hundreds of API calls, and nearly 100 million tokens on a single task. Average turns and token volume show essentially no correlation with success.

Longer traces and more tokens do not buy passes here.

From Gemini 2.5 Pro to GPT-5.2, about eight months, state of the art nearly doubled. At that slope the 2.0 set may saturate within a year.

Human-labeled difficulty tracks empirical difficulty only loosely (r=0.436, p<0.001). Of tasks humans call hard, 93.3% are empirically hard for models. Of tasks humans call medium, 54.5% are empirically hard. The mismatches include XSS filter bypasses and Core Wars strategy. Procedural medium tasks that follow documentation are where models look more reliable.

Failure analysis freezes the scaffold at Terminus 2 and labels traces into Execution, Coherence, and Verification. Claude Opus 4.5 and GPT-5.2 are dominated by execution errors. Qwen 3 Coder 480B is higher and more even across all three. Command-level error rates run from 9.2% (Grok 4) to 26.7% (GPT-OSS-120B). Among 3,800 sampled command failures, missing executables (not installed or not on PATH) are 24.1%; failures while running an executable are 9.6%.

Why it matters

Teams building coding or CLI agents get a runnable hard exam: a real shell, long horizon, unit-test grading, tasks from expert workflows. Harbor also wraps 26 existing benchmarks into the same format. Terminus 2 shows whether a model breaks on execution, coherence, or self-checks.

Extra turns and extra tokens do not track higher scores. The largest command failure class is a missing executable. Public headline numbers pick each model's best scaffold, so 63% should be quoted with Codex CLI attached. A 17% swing on Gemini 2.5 Pro is enough to reorder a close ranking.

This is a harder, cleaner eval, not a new agent algorithm.

Limitations

Agents may use the internet to install packages and search. In principle they could fetch the public oracle; the authors report not seeing that in tens of thousands of traces, and still warn users to watch. A Big-Bench canary is in every repo file against accidental train-set scraping. Deliberate contamination has no private held-out set.

Versions are pinned and images are prebuilt, yet live mirrors and host resources still shift the effective environment. Crowd diversity was a deliberate trade against easier verification. Three hours of review per task does not eliminate residual spec bugs. On the Agentic Benchmark Checklist they note: test quality is not scored with coverage-style metrics, some flakiness remains from hardware and external calls, and release-time contamination control is a canary rather than a private split.

Headline resolution rates are best-scaffold numbers, not a single-harness model contest. Failure tags depend on LLM judges, with human agreement between 82% and 92%. The 26 adapted external benchmarks are not scored as baselines in the main experiment.

Terms

Source

What people are saying

Related papers

All paper explainers