NTU Study: Switching Agent Harness Can Reverse Claude vs GPT Rankings

机器之心 · wechat · 2026-10-04

A NTU team led by Prof. Bo An systematically cross-tested 4 agent harnesses (OpenHands, DSH, PI, openJiuwen) × 5 models × 3 task sets, plus Codex–GPT and Claude Code–Claude native pairings. Key findings: harness choice can flip model rankings (Claude Opus 5 leads GPT by 8 points on Terminal-Bench4 with OpenHands, but loses by 30 with PI); 4 of 5 models change best harness by task, while openJiuwen–Kimi K3 wins consistently thanks to split tool interfaces, timeouts, and truncation-resume; native pairings aren't optimal; and cost ≠ quality — DSH spent 4.3× PI's $293 on GPT with worse scores, and openJiuwen hits 94–99% prompt-cache rates. Conclusion: evaluate models and harnesses jointly on real tasks.

Original post →

More from coding & agent

coding & agent channel →