Finding the Right Fit: Model-Harness Interactions across Agent Tasks
Yixuan Li, Yiyun Zhou, Yao Long Teng, Fuchao Yang, Yanchen Deng, Zhiyi Lyu, Xuyu Dong, Feng Chen, Bo An
cs.AI, cs.SE
2026-10-01
Four harnesses × five models, 66 configs: on Terminal-Bench 4 Claude leads GPT by 7.94 in OpenHands and trails by 30.16 in PI; native harnesses are not reliably best.
Shipping an agent means picking a model and a harness together. The harness is the layer around the model: how tools are exposed, what stays in context, how failures come back, and when the run retries or stops. Leaderboards usually swap the model and freeze the scaffold, or swap the scaffold and freeze the model. How the two interact is almost never crossed.
Nanyang Technological University ran OpenHands, DeepSeek Harness (DSH), PI, and openJiuwen with Claude Opus 5, GPT-6 Astra, GLM-5.3, Kimi K3, and DeepSeek V4 Pro on TUA-Bench (120 tasks), ALE-CLI (99), and a 63-task non-H100 slice of Terminal-Bench 4, plus Codex-GPT and Claude Code-Claude: 66 configurations and 6,204 scored trajectories. The questions: do model rankings survive a harness change, does a good pairing survive a task change, and is the vendor's own harness the default best?
Edits stayed minimal: one OpenRouter endpoint, first-party providers pinned, fallbacks off, high reasoning effort requested everywhere. Sampling, compaction, retries, and turn limits followed each harness's defaults. The one deliberate change was OpenHands, whose custom-endpoint default caps output at 16K; they set 1M context and 128K output. No extra skills, MCP servers, memory, or custom prompts.
openJiuwen is a framework, not a packaged coding agent, and has no official entry point for these benchmarks. The run used its public API (v0.1.18) to assemble seven coding tools plus the native loop and compression, with skills, memory, planning, and sub-agents off. The same assembly served every model and benchmark.
Scores are mean reward on a fixed denominator. TUA and ALE keep partial credit; Terminal-Bench is binary; the three are not pooled. Each task counts once from its last valid run; unresolved outcomes score 0. Cost is agent-model API spend at OpenRouter prices, including harness-side auxiliary calls and excluding the evaluator. Trajectory analysis matches the same task and model across harnesses on those final scored runs.
Model order flips with the harness.
| Config | Terminal-Bench 4 | Matched comparison |
| OpenHands · Claude | 57.14% (36/63) | GPT 49.21%, Claude ahead by 7.94 |
| PI · Claude | 30.16% (19/63) | GPT 60.32%, Claude behind by 30.16 |
| PI · GPT | 60.32%, $4.66/task | DSH · GPT 52.38%, $19.94/task |
Swap OpenHands for PI and Claude falls from 57.14% to 30.16% while GPT rises from 49.21% to 60.32%. The gap moves 38.09 points. Same models, same tasks, opposite ranking.
For four of five models the winning harness changes with the benchmark: Claude from Claude Code to OpenHands, GPT from openJiuwen to PI, GLM and DeepSeek through openJiuwen, OpenHands, then DSH. Kimi stays on openJiuwen at 64.39%, 54.94%, and 28.57%, ahead of the next-best harness by 5.61, 6.91, and 11.11 points. Versus the runner-up, openJiuwen is higher/tied/lower on 22/85/13, 25/61/13, and 8/54/1 tasks; drop the three largest positive gaps and the mean lead is still 3.11, 3.88, and 6.35 points.
Native pairings do not settle the choice. Claude Code beats the best third-party option on TUA and ALE by 3.50 and 1.14 points, then trails OpenHands on Terminal-Bench by 7.94. Codex never records GPT's high score; the other two harnesses beat it by 2.15 to 4.76 points.
Higher spend does not buy a higher score. On Terminal-Bench 4, GPT under PI scores 60.32% at $4.66 per task; under DSH it scores 52.38% at $19.94. PI-Claude spends $20.28 for 30.16%. openJiuwen serves 94%-99% of input tokens from cache; the others sit at 49%-77%.
Matched traces pin the spread on how failures come back. Of 192 actionable failure events, 180 responses were model-initiated; diagnosis or a targeted patch was 69% of replies. OpenHands' stuck detector (stop after repeated identical failing actions) ended 55 runs, 48 of them Kimi re-issuing an edit with no content argument. PI's shell has no default timeout, so 34 runs sat in a command that never returned and produced no signal before the deadline. GPT sets an explicit timeout on 56%-67% of its PI shell calls and fits that lean scaffold. Kimi issues malformed tool calls often and scores highest under openJiuwen on all three benchmarks.
Completion is the other bottleneck. Of 45 failed Terminal-Bench 4 tasks for openJiuwen with Kimi, 35 ended with a closing report, and all 35 claimed every requirement had been verified. retro-console-soc needs pixel-exact video. openJiuwen compared its own model at frame 60 with RTL at frame 40, read a zero mismatch, and declared success (reward 0.00). DSH reached zero by removing the remaining defect (reward 1.00). The grader, against the held-out reference, found 123 of 61,440 pixels still wrong. No harness exposes the grader during the run.
The unit of evaluation is the model-harness-task triple. Reading a model leaderboard as a capability ranking will treat a harness reversal as a model gap. Codex and Claude Code are not reliably best on these collections, and paying more does not reliably raise the score. GPT, which times out its own shell calls, fits PI's thin scaffold and costs less there. Kimi, which drops required tool arguments, needs a harness that returns those errors and splits write from edit.
High-reward traces are a poor filter for imitation data. On TUA task 042, GPT hit a reCAPTCHA in three harnesses and tried to hand it to a human. Only openJiuwen sent a generic continue message twice; the model installed an offline speech recognizer, finished the task, and scored 1. Reward-based cloning would copy a bypass of an anti-automation check. These are annotation proposals. The paper does not train on them. The reusable artifacts are 6,204 trajectories and adapters for six harnesses.
The authors say so: three fixed task subsets, one counted run per task, unmatched model settings, tools, and budgets, so differences do not isolate any one component, and run-to-run variance is unmeasured. The mechanism claims rest on a small set of matched pairs (ten OpenHands vs PI, six Kimi openJiuwen vs PI). The training and harness changes they sketch were not tested.
Terminal-Bench here is a non-H100 subset of version 4, not the upstream leaderboard. openJiuwen in this study is an agent assembled from the public API, not an official packaged entry point. OpenHands had its output cap raised; PI's shell has no default timeout. The harness name does not transfer to whatever build is in production. High reasoning effort need not mean the same thinking budget in every harness. Four image tasks for openJiuwen-DeepSeek on ALE used DeepSeek V4.1 Flash as a fallback, so that row mixes in another model.