NTU Study: Agent Harness Quality Depends on Model-Task Combo, No Universal Winner

jiqizhixin · x · 2026-10-11

A study from Nanyang Technological University (An Bo's team), "Finding the Right Fit: Model–Harness Interactions across Agent Tasks," systematically evaluates OpenHands, DSH, PI, and openJiuwen harnesses with Codex–GPT and Claude Code–Claude native pairings as references. Key finding: harness quality is not a fixed property but depends on the specific combination of harness, model, and task — tool calling, context management, and error recovery all shift across combinations, with no universal winner. Vendor-native harnesses are not always the best fit for their own models.

Original post →

More from coding & agent

coding & agent channel →