Automated agent harness search vs human taste: no winner across drug design tasks

niloofar_mire · x · 2026-10-03

A new blog asks whether human taste is overrated in agent harness engineering — the harness being the system around a model that decides tool calls, evidence, and memory across steps. The authors ran Qwen on three drug design tasks from SMDD-Bench, comparing a manually redesigned harness against one produced by automated harness search with Claude in the loop.

Key findings:

The piece reflects the broader shift of harness iteration being automated by strong models rewriting the harness itself — relevant to anyone doing agent engineering.

Related event: CMU Blog: Automated Harness Search vs Manual Tuning for Science Agents(2 posts)→

Original post →

More from coding & agent

coding & agent channel →