NTU Study: Switching Agent Harness Can Reverse Claude vs GPT Rankings
机器之心 · wechat · 2026-10-04
A NTU team led by Prof. Bo An systematically cross-tested 4 agent harnesses (OpenHands, DSH, PI, openJiuwen) × 5 models × 3 task sets, plus Codex–GPT and Claude Code–Claude native pairings. Key findings: harness choice can flip model rankings (Claude Opus 5 leads GPT by 8 points on Terminal-Bench4 with OpenHands, but loses by 30 with PI); 4 of 5 models change best harness by task, while openJiuwen–Kimi K3 wins consistently thanks to split tool interfaces, timeouts, and truncation-resume; native pairings aren't optimal; and cost ≠ quality — DSH spent 4.3× PI's $293 on GPT with worse scores, and openJiuwen hits 94–99% prompt-cache rates. Conclusion: evaluate models and harnesses jointly on real tasks.
More from coding & agent
- Rogue agentic deployments should be tracked as APTs, researcher argues — nitarshan · 2026-10-11
- Open Dots: MIT-licensed self-hosted agent workspace with 5.6k GitHub stars — matchaman11 · 2026-10-11
- From 'Vibe Coding Is Evil' to 'My New Game Is All Vibe-Coded' in 10 Months — DeryaTR_ · 2026-10-11
- NYC AI dinner leak claims all frontier labs run "RSI loops" (unverified) — Hesamation · 2026-10-11
- Garry Tan's multi-agent workflow: one thread runs the merge PR queue, only green CI hits master — garrytan · 2026-10-11
- Plane Agents Burn Massive Tokens Just Two Weeks After Launch — JosephJacks_ · 2026-10-11