Harness design, not the model, drives coding agent scores: 176-setup study quantifies it

alex_verem · x · 2026-09-25

Researchers from UMass Amherst, Zoom, Emory, and UNC Charlotte held models constant and varied only the harness—planning, action space, context management—across 176 matched setups on SWE-Bench Verified and Terminal-Bench 2.1.

Key findings:

Conclusion: no single best setup—pick components per model, task, and budget. A benchmark score measures model plus harness, and the harness controls a large share of it.

Related event: 176 Controlled Experiments Quantify How Agent Harness Design Affects Performance(2 posts)→

Original post →

More from coding & agent

coding & agent channel →