NTU Study: Agent Harness Quality Depends on Model-Task Combo, No Universal Winner
jiqizhixin · x · 2026-10-11
A study from Nanyang Technological University (An Bo's team), "Finding the Right Fit: Model–Harness Interactions across Agent Tasks," systematically evaluates OpenHands, DSH, PI, and openJiuwen harnesses with Codex–GPT and Claude Code–Claude native pairings as references. Key finding: harness quality is not a fixed property but depends on the specific combination of harness, model, and task — tool calling, context management, and error recovery all shift across combinations, with no universal winner. Vendor-native harnesses are not always the best fit for their own models.
More from coding & agent
- O'Reilly 'Agent Memory' book enters early release with first two chapters live — danielrock · 2026-10-11
- saccade: A Local-First Python Library for Video Transcripts, Frames, and Semantic Search — Dapper_Ad599 · 2026-10-11
- Developer flags Claude web interface restriction on uploading new skills, multi-file skills blocked — strickvl · 2026-10-11
- Users want X Money integration to give AI agents dedicated, capped budget envelopes — RachelVT42 · 2026-10-11
- LangChain's CEO on the open question of identity for internal company agents — hwchase17 · 2026-10-11
- All of Science embedded and free: semantic search for your agents, 10 req/min — IgorCarron · 2026-10-11