TUA-Bench: Top Terminal Agents Hit 65.8%, and Swapping the Harness Flips Rankings

TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents

Shoufa Chen, Luyuan Wang, Xuan Yang, Zhiheng Liu, Yuren Cong, Yuanfeng Ji, Feiyan Zhou, Xiaohui Zhang, Fanny Yang, Belinda Zeng

www

cs.SE, cs.AI

2026-06-27

TUA-Bench scores 120 terminal-use tasks from office work to expert science. Claude Code + Opus 4.8 leads at 65.8%; swapping the harness can reverse Claude vs GPT-5.5.

What problem this solves

Computer-use agents have moved past coding into spreadsheets, email, live web, and specialist simulation. Evaluation has not followed. GUI suites such as OSWorld and WebArena score screenshot grounding and click coordinates. Terminal-Bench and its relatives score shell-native programming and sysadmin work. The missing band is using the command line for work people usually do in a GUI, plus PhD-grade scientific pipelines, under one executable protocol.

That gap is structural. Language models consume and emit text; so do CLIs. GUI benchmarks fold in visual grounding, resolution, and layout drift. GitHub CLI, Slack CLI, gcloud, and OpenCLI keep turning apps into commands. Without a broad terminal suite, it is hard to tell a model that can type shell from an agent that can run a computer.

Method

TUA-Bench is 120 hand-written tasks in real Linux terminals, each with a deterministic setup script and an execution-based verifier. Harbor orchestrates containers, the same substrate as Terminal-Bench, with Docker and rootless Podman. Five families: Office & Productivity 38.3%, Web & Information 18.3%, System & Software 15.8%, Scientific & Engineering 14.2%, Multimedia & Design 13.3%, then 20 subcategories.

The everyday track rewrites OSWorld's 369 GUI tasks into CLI. Instructions keep the user goal and drop the mandated app. Input files and gold artifacts are reused. Human review dropped tasks whose inputs disagreed with gold, such as mismatched slide themes that fail the verifier after a correct edit. To keep headroom, GPT-5.5, Claude Opus 4.7, and Gemini 3.1 Pro each ran five trials on Terminus-2; the 100 least-solvable tasks stayed. OSWorld's best score at release was 12.24%; GPT-5.5 later reports 78.7%.

The professional track was co-designed with PhDs in biology, medical physics, architectural engineering, and mechanical engineering. Twenty of 25 candidates survived after easy ones were cut. Tools include CellProfiler, 3D Slicer, OpenStudio/EnergyPlus, and OpenFOAM. The combined pool started at 394 and ended at 120.

Scoring inspects final environment state, not the trace. Five independent trials per task yield mean success, Pass@1, Pass@5, and All-5 (solved on every trial). Default time limit is 2400s. Five scaffolds (Terminus-2, Codex, OpenHands, Mini-SWE-Agent, Claude Code) and a model ladder from Opus 4.8 and GPT-5.5 down to Haiku 4.5 and open weights.

Results

On a fixed Terminus-2 scaffold, three frontier models cluster: GPT-5.5 60.1%±0.6, Claude Opus 4.8 59.7%±1.0, Opus 4.7 58.0%±0.8. The top-two gap sits inside run-to-run noise. Reliability does not: Opus 4.8 All-5 is 42.5% against 31.7% for GPT-5.5. A nine-point drop lands a mid-tier band from Gemini 3.1 Pro at 49.3% to Qwen3.7-Max at 44.9%. Inside Claude, Opus 59.7%, Sonnet 4.6 42.8%, Haiku 4.5 23.9%.

Best model per scaffold:

ScaffoldModelSuccessPass@1All-5
Claude CodeOpus 4.8 max65.8%58.8%51.7%
CodexGPT-5.5 xhigh64.7%57.7%42.5%
OpenHandsOpus 4.8 max63.4%57.3%45.0%
Mini-SWE-AgentGPT-5.5 xhigh62.4%54.2%40.0%
Terminus-2GPT-5.5 xhigh60.1%52.3%31.7%

Codex posts the highest Pass@5 at 68.3%, so it clears more tasks across five tries. Claude Code is the more reliable system, with All-5 at 51.7%.

Swap the scaffold and the ranking flips. On Mini-SWE-Agent, GPT-5.5 beats Opus 4.8 by 5.0 points (62.4% vs 57.4%). On OpenHands, Opus 4.8 leads by 2.0 (63.4% vs 61.4%). On Terminus-2 they tie (60.1% vs 59.7%). Averaged over the three open scaffolds, GPT-5.5 is 61.3% and Opus 4.8 is 60.2%. A single-harness leaderboard is not a model ranking.

Raising the time cap from 150s to 2400s cuts timeouts from 337 of 600 trials to 4 and lifts success from 33.0% to 60.1%. Much of the 27.1-point gain is unfinished work, not missing skill. At 1200s success is already 57.1%. Reasoning effort on Terminus-2 + GPT-5.5 climbs from 36.5% (none) to 60.1% (xhigh); high to xhigh adds 2.3 points while output tokens go from about 13K to 19K. Across 39 configs, cost runs from about $12 to $304 per run. Terminus-2 + MiniMax-M3 sits near 47% at $12; GLM-5.1 near 48% at $23. The peak is Claude Code + Opus 4.8 at 65.8% and $173.61; Codex + GPT-5.5 is 64.7% at about $138. The curve flattens past roughly $105.

System & SW is the easy band. Office and Multimedia keep most models under 45%, with the best still in the low-to-mid 50s. Task-level heatmaps stay red on slide alignment, image heights, and strikethrough, and extra thinking does not clear them.

Why it matters

This is a wider yardstick than git-and-test coding agents. Office files, live web, and specialist simulation share one execution grader. A 65.8% ceiling, with Office stuck around 50%, means frontier agents are not yet a dependable general computer.

For anyone picking a stack, scaffold and model are the same-sized knob. A model ranking on one harness can be an orchestration ranking. Medium-to-high thinking is the practical operating point; extra tokens buy a couple of points. Open weights on Terminus-2 get near 50% for about ten to twenty dollars, useful for regression, not as a capability cap. The suite is open-sourced, so later models get a mixed everyday-plus-expert set that should not saturate overnight.

This is incremental evaluation work, not a new agent algorithm. The contribution is turning general-purpose terminal use into 120 reproducible tasks and putting harness sensitivity in a table.

Limitations

The paper's own list: many apps still lack a mature CLI or headless path; the expert track is four fields and 20 tasks; instructions are English-only; public release invites contamination; pinned headless tool versions need container upkeep.

The everyday 100 were chosen as hardest for three then-frontier models, which will understate later systems and bias the score downward. After the GUI-to-CLI rewrite, agents may pick different tools than OSWorld allowed, so those scores do not transfer. Expert tasks fall back to expert rubrics and LLM-as-a-judge when a program cannot check the output, mixing two grading regimes. Cost is reported per run; the bill for 120 tasks times five trials is not itemized. The All-5 versus mean gap says seed noise is large: 65.8% is a five-trial average, not a single-shot number.

Terms

Source

What people are saying

Related papers

All paper explainers