Automated agent harness search vs human taste: no winner across drug design tasks
niloofar_mire · x · 2026-10-03
A new blog asks whether human taste is overrated in agent harness engineering — the harness being the system around a model that decides tool calls, evidence, and memory across steps. The authors ran Qwen on three drug design tasks from SMDD-Bench, comparing a manually redesigned harness against one produced by automated harness search with Claude in the loop.
Key findings:
- No single approach won everywhere. The manual harness pulled clearly ahead when the fix was changing what the agent sees and remembers.
- Once that was in place, automated search did better when the remaining problem was how to search the chemistry itself.
The piece reflects the broader shift of harness iteration being automated by strong models rewriting the harness itself — relevant to anyone doing agent engineering.
Related event: CMU Blog: Automated Harness Search vs Manual Tuning for Science Agents(2 posts)→
More from coding & agent
- New Claude Code plugin renders pasted images as thumbnails above the prompt — jarrodwatts · 2026-10-03
- Coinbase for Agents adds bracket, stop-limit and TWAP orders plus automated feedback loop — kleffew94 · 2026-10-03
- PhantomEnvironments: 7B LLM Trained in Synthetic RL Environments Matches Agents 10x Its Size — CShorten30 · 2026-10-03
- OpenAI introduces dot: cross-app memory, context and Codex task coordination — OpenAIDevs · 2026-10-03
- MIT Interactive Diagrams: From Attention to Mixtral and DeepSeek-V3 Architectures — vtabbott_ · 2026-10-03
- Agent Aware Architecture: an open spec to make sites walkable for AI agents — sierracatalina · 2026-10-03